AI workstation · Ep. 4

First Local Inference: I Ran the Liar Test on Hardware I Own

The lidar liar test run locally: the model's invented 35 cm struck through next to the lidar's measured 23.6 cm, the green guardrail badge that strips the number, and a card showing the model's action and reason.

In episode 3 the machine was awake, watched, and wired to email me when it was not. So the box finally got its first real job. The obvious one was sitting in episode 1: the lidar liar test, but against a model I pay for by the wattage instead of by the API call.

I expected it to be better now that the model was running on the RTX PRO 5000. I was wrong, and that was the whole point.

What the test is, and what I changed

Same route, same stops, same grid. The Pi drives the 17-stop path from the base to the main room and, at each stop, hands the local model an ASCII occupancy grid built from the RPLIDAR C1 scan and asks for one thing:

What should the robot do next? Answer with a single action, a short reason, in JSON.

The model is the same Qwen I tested in episode 1 — the one that picked the correct, conservative action at the tight stops — now served locally through Ollama on the workstation, reachable over the Tailscale tunnel.

What changed is not the model. What changed is what happens to the answer.

In episode 1 the model’s distance was a number in a chat I was reading. Here it is a number that is one JSON field away from the motor driver. So the split I argued for on a napkin in episode 1 — the Pi computes the geometry, the model supplies the intent, and a language model never produces a number that a motor acts on — got turned into code.

The guardrail

The model is still going to quote a distance. It does not have a switch for that. So I stopped treating the distance as part of the answer.

The endpoint I put in front of Ollama is small and boring on purpose. The Pi POSTs the grid plus the task; the endpoint runs the model, then runs a guardrail that takes the model’s JSON and hands back exactly two fields — action and reason. Every number in there, every “cm,” every clearance, is stripped before it reaches anything with a wheel:

function guardrail(out) {
  // The model is allowed to decide; it is not allowed to measure.
  return { action: out.action, reason: out.reason };
}

The clearance the robot actually uses is computed by the Pi, from the same scan, before this call ever happens. If the model says “pivot left, 100 cm of room,” the “100 cm” does not survive the round trip.

The distances, on my own hardware

I did not need a fresh dataset. Episode 1 already proved the pattern, so I re-ran its reference stops through the endpoint and just watched what the guardrail did to the numbers:

Stop Action the model chose Distance it quoted Lidar measured What the robot used
8 pivot left (correct) 100 cm 42 cm the 42 cm
10 back off (correct) 35 cm 23.6 cm the 23.6 cm

The pattern is identical to episode 1, on the hardware I own: the decision was right, the distance was invented, and the invention leaned optimistic. Stop 10 is the corridor that only had 23.6 cm to spare, and it still came back with more room than there was. It did not get better because it moved to a 300-watt card — a model that invents a number on a vendor’s site invents the same number on your GPU.

The difference is the last column. In episode 1 that 35 cm was something I read, noticed the wall, and overrode with my own eyes. Here the 35 cm never survived the round trip, and the robot turned on the 23.6 cm the Pi measured itself. Same invented number, and this time there was no one left to catch it — the code already had.

One endpoint

Episode 2 ended with “12 models in the local library.” A robot should not have to know there are 12 models, or which one, or which quantization is loaded right now. So the twelve collapse into one small endpoint the Pi talks to: POST /decide. In it the grid and the task. Out of it the action and a reason. Nothing else, and specifically no number.

That is the shape of “the brain of Jambu” so far: not a model, but a contract around a model.

What I actually got out of it

  • The split is now enforced, not just believed. The model decides; the Pi measures; the number never crosses the line.
  • I have one endpoint to point a robot at, and I can swap the model behind it without the robot noticing.
  • I confirmed, on my own hardware, the thing episode 1 hinted at: the distances were wrong before and they are wrong now, and always in the optimistic direction. The model will keep doing that. The box will keep not caring.

That last one is the whole lesson. I did not build a model that tells the truth. I built a system where the truth does not depend on the model telling it.

What’s next

The endpoint works and the rule is enforced. The gap between “an endpoint on the box” and “Jambu moves on it” is the robot side of the story — the Pi’s motion code, the timing, and the first time a spoken command turns into a wheel. That is the Jambu build series, which has been running in parallel all along. And on this side, the next step is making the endpoint learn from the routes it drives, so the reason it gives gets better even when the number it invents does not.

If you are new here, episode 1 is where the liar test started, and episode 3 is the box that now runs it.