"Come here"
He hears "come here," turns toward my voice, finds my face across the room, and drives to me. Forty-five seconds. It took three days to earn.
Seven systems, one sentence
"Come here" sounds like one command. It is actually seven systems that have to be telling the truth at the same moment:
- The mic array measures which direction the voice came from.
- The tracks turn that many degrees — and have to know they did.
- The camera has to find a face at four feet.
- The lidar has to gate the path so he doesn't drive through a bar stool.
- The encoders have to notice if he's pushing furniture instead of moving.
- The arrival rule has to decide, honestly, whether the thing in front of him is a person or a chair.
- And the head has to be pointed where a face would actually be.
Every one of those lied to me over three days.
The liars, in order
The compass was 90° off. The mapping software and the lidar disagreed about which direction counted as forward — and then, when I corrected it, the map came back exactly backwards, which turned out to be the fingerprint of a mirrored angle convention between the sensor and the SLAM library. Two calibration constants, opposite signs, both correct.
The ears chased his own neck. The mic array localizes a real voice beautifully — you can watch its LED point right at you. Then the head servo moves, the whine becomes the loudest sound in the room, and the bearing follows the servo. Same problem with his own text-to-speech: the loudest speaker in the room is bolted to his chest. The fix is that he goes deaf to himself — every time he talks, turns his head, or rolls his tracks, the direction-finder stops accumulating.
The camera was the wrong camera. A 4K webcam exposes several video devices, and the browser had quietly picked the infrared one at 640×480. He'd been seeing the world through his night-vision eye for weeks. That single default explains everything downstream: faces only detectable at arm's length, blindness across a room, and every resolution fix I tried "not working" — I was tuning a camera nobody was looking through.
The vision pipe wasn't connected. Frames went from the kiosk to the robot's local server, and nowhere else. His conversational brain — the part that answers "what do you see?" — had never once received an image. So it answered from memory: a workbench in a room he hadn't been in for months. He wasn't hallucinating. He was reminiscing, because nobody had ever handed him a photograph.
And he was looking at feet. After everything else was fixed, faces still vanished. He's not real tall, so his head calibration had him aimed about 30 degrees low, right at my feet, with my face cropped off the top of the frame. The detector was right to report nothing; there were no faces in the picture. Recalibrating the head servo to true horizontal brought the faces back.
How you actually find these
Not by reading code. Every one of those bugs was found the same way: point the robot at something known, look at what the sensor actually reports, and believe the floor over the theory.
- "The arrow is off 90 degrees right" → a mounting calibration.
- "Down a few degrees… boom" → the horizon, by eye.
- Speak from dead ahead, then from his right, read two numbers → the mic's zero and its rotation sense.
- Pull the frame the camera is currently holding and just look at it → he was looking at my feet. Recalibrate to horizontal.
The robot cannot tell you it is confused. It can only act confused. The instrument that finds these bugs is a person standing in the room saying "no, he's looking at my feet."
What it does when it works
Voice bearing first. If the face isn't found in the first sweep, he rotates and sweeps again, and creeps toward the voice while looking — distance beats resolution eventually. When he finds you, he says "I see you", because when it works you want to hear it. If he never finds a face, he doesn't pretend: he comes toward the voice and asks where you are.
He arrives at about 45 centimetres, then looks up, because he is short and you are not.
The stack
Everything in the video runs on the robot — no cloud, no network required:
- Jetson Orin Nano — vision, mapping, voice, and all the reflex loops
- ReSpeaker XVF3800 — 4-mic array, direction of arrival
- Logitech BRIO — the eye, owned by a native sidecar (no browser)
- LDRobot LD19 — 360° lidar for mapping and obstacle gating
- ESP32-S3 "track brain" — motor control over Bluetooth, so the body works in buildings whose WiFi I don't control
- Two down-facing ToF sensors as cliff eyes, because a 2D lidar cannot see a drop-off
Full bill of materials, wiring, pin map, and all 28 printed parts: the v3 build post. STLs: meckie-all-parts.zip.
What's next
The map now completes itself — he drives to the edges of what he knows until there are no edges left. Named places work ("this is the kitchen"), and routes plan through the floor plan he built. Which means the next video is him going somewhere on his own.