New jobs from skills the robot already has
“After a storm, check the back gutter.” Kith plans with the robot's own skills (go to, look, compare) and checks the result.
How Kith works

At the top, a general vision-language model (VLM) does the thinking. It reads instructions, understands camera images, and plans the job. It is too slow to control motors, so it is never in the real-time loop.
In the middle, Kith turns the plan into a program of small steps and runs it. Anything that must happen fast, like spotting a ball in flight, is handled by small, fast parts that Kith builds for the job.
At the bottom, the robot's own controller, or a robot action model, moves the motors and keeps the robot safe.
Kith checks each step with the robot's cameras and sensors. It learns by practice and from people's advice, and saves what it learns as programs and notes that people can read and correct. The large AI model is never retrained.
“After a storm, check the back gutter.” Kith plans with the robot's own skills (go to, look, compare) and checks the result.
“Catch without a bounce: match the ball's speed.” Practice fine-tunes the motion, and a coach's advice can change the motion itself. No AI model is trained.
What Kith learns is kept as plain text you can read, not hidden inside an AI model.
Kith is software, and it can work with a robot in more than one way. It can run on the robot itself, if the robot has a suitable processor. Or it can run on a computer nearby and control a simpler robot, such as a basic drone, from there. It can work with smart robots that take a task in words, and with simple ones that take only speed and turn. Just how Kith connects to each robot, we will work out with our design partners.
Whatever the setup, your robot's own safety system always has the last word. Kith can send a stop request at any time, and your robot decides how to stop safely.
Kith builds these tools by itself, on the fly, to suit each robot and each job. No AI or software expertise is required on your part.
A small vision model that sees one thing very fast, where a big AI model would be too slow: a ball in flight, a part on a belt, a defect. It is trained automatically, with no labeling by hand; a general AI model labels the images and keeps checking its work.
It practices a motor skill on your robot, or on its simulator, and adjusts the skill's settings after each round. It measures every miss, learns how your robot really moves, and gives each new skill a head start.
In a busy or very large scene, Wisdom Eye makes zero-shot or one-shot object recognition much faster than a standard vision-language model can manage alone.
Kith comes with modules. Each module is a set of ready-made skills, and a maker picks the modules for a robot. Some modules already work in simulation; others are planned.
| Module | What it gives your robot |
|---|---|
| Sentinel Kith's own semantic SLAM | Maps a site quickly, recognizes places and things, goes where told in words (“the room past the kitchen”), patrols, takes stock, and reports what changed as the owners wish |
| Swarm | Organizes many robots and drones of different capabilities into teams. It hands out the work, keeps the teams in step, covers for any robot that fails, runs scheduled patrols, upgrades its robots on the fly with Mind Patches, and raises the alarm by the site's own rules |
| Perception | Learns to recognize a new part or product from a single example picture, and sharpens what a robot's small on-board AI model can see with a Mind Patch, with no retraining |
| Motion skills | Toss, catch and gentle catch, learned by practice and carried to a new arm |
| Sorting station | Sorts mixed objects by rules a person states and changes, handling each object with the care it needs |
| Human skills | Learns the people a robot works with: their authority level, preferences, intentions and groups |
| Teacher | Once Kith has mastered a job, teaches it to people |
Watch two simulated robot arms go from jerky first tries to smooth, accurate juggling.