AI workstation · Ep. 5

Teaching the Endpoint to Learn From Its Own Routes

The offline replay loop — replay, judge, fix the reason — next to a stop-10 card where the weak reason is struck through and the fixed reason earns a higher judge score.

Episode 4 ended with the rule enforced: the model decides, the Pi measures, and no number the model invented ever crosses the line. It also ended with a promise — making the endpoint learn from the routes it drives, “so the reason it gives gets better even when the number it invents does not.”

Here is where that promise had to be renegotiated. The endpoint does not learn. It doesn’t have anywhere to put what it learned. And after staring at that for a while, I realized that was the correct behavior, and that the thing I actually wanted was something I can build.

The endpoint can’t learn, and that is the point

POST /decide is stateless on purpose. In: the grid and the task. Out: an action and a reason. There is no memory behind it. Two runs with the same input give the same shape of answer, and nothing the robot does afterwards is written back anywhere the model can read.

That is not a limitation I failed to fix. That is the guardrail of episode 4 at the model level: if the endpoint updated itself from the robot’s history, then a bad run — a wrong action, a lucky one, a sensor glitch — would become training signal, and “a language model never produces a number a motor acts on” would quietly become “a language model that the motor’s own history shaped.” The contract in episode 4 was the whole win. I am not going to eat it in episode 5.

So the endpoint stays a pure function of the input. Which means “learning from routes” has to live somewhere the motor never sees.

Replay: the route is the teacher

Every route Jambu has driven is already saved — episode 1 kept the lidar scan, the ASCII grid, the photo and the action the robot actually took, at every stop. That is a teaching set with ground truth I trust, because the ground truth is what a real robot did with real wheels, not what a model claimed.

So the loop is offline, and it never touches the motor:

  1. Take one saved route — a stop with the grid, the action that was actually taken, and the outcome.
  2. Re-run the endpoint on that exact grid. Same model, same prompt. It comes back with an action and a reason.
  3. Have a judge grade the reason only: does this reason actually explain why this action is right for this grid?

The judge is the same model family, run as a second call — and it grades a string against a saved route, not a number against a wall. There is no motor behind it. There is nothing to hit.

// Offline. No motor, no wheels, no risk. Just two calls.
function replay(stop) {
  const answer = endpoint("/decide", stop.grid, stop.task);
  const grade  = endpoint("/judge", { route: stop, answer });
  return { action: answer.action, reason: answer.reason, score: grade.score };
}

If the reason is weak — “low risk” about a corridor that had 23.6 cm to spare — I fix the reason text I ship with the prompt, or the way I ask, and replay until the grade holds. The model’s weights never move. I move the words around the model.

What the grades actually looked like

I replayed episode 1’s route three times, one round of fixing between each pass. Same model, same grid; only the wording I control changes.

Stop Pass 1 Pass 2 Pass 3 What the judge kept flagging
3 5.5 8.5 8.5 “the wall” named, not located
8 4.0 7.5 9.0 gap width still unconfirmed in words
10 6.5 9.0 9.0 —

Stop 8 is the one that tells the story. Pass 1 came back with “narrow gap, low confidence” — the exact kind of confident vagueness that looked safe in a chat window in episode 1. The judge gave it 4.0, because “narrow gap” names nothing. By pass 3 the reason said where the gap was and why the pivot was the move, and it held at 9.0 across two re-runs.

The distance it still invents is the same distance it always invented. That part did not change, and I stopped expecting it to.

What the judge can’t tell you

Two honest limits, both discovered the hard way.

The judge is also a language model. It, too, invents. It grades “confident but vague” reasons differently on two identical runs, and it invented its own clearance figure once — 60 cm — in a note about a corridor that had 23.6. So I treat its score as a ranking and a trend, not a measurement, and I only trust a grade that survives a re-run. A judge that can be wrong is still better than no judge, but it is a second model, not a truth source. The truth source is still the lidar, as it has been since episode 1.

The score is not the number. This is the part worth saying plainly. The whole point of the loop is that “the model told the truth” is not a reachable goal — episode 4 proved that on hardware I own — so I am not chasing a grade of 10. I am chasing a reason that explains the action, that would still make sense to a person reading the log at 2 a.m., and that the robot’s real history supports. The number is wrong. The reason stopped being.

What I actually got out of it

  • A loop that improves the endpoint’s output without the endpoint changing. The contract from episode 4 is intact: stateless in, action and reason out, no number.
  • A way to say “it got better” with a record behind it — the same route, re-run, graded — instead of a feeling.
  • A confirmed division of labor at the model level: the box learns nothing, and I don’t need it to. It needs to be right and explainable for the grid it was given. The learning is mine, offline, where a mistake can’t be driving anything.

That last one is the lesson, and it is the same lesson as episode 4, one layer up. I did not build an endpoint that learns. I built a loop that makes the words around a non-learning endpoint better, and I can show you the route that taught me.

What’s next

The reason now survives a re-run, and the loop is cheap to point at any new route. The gap that’s left is the one that has run in parallel the whole time — the Pi’s motion side, where an action becomes timing and a wheel. That is the jambu build series. And on this side, the next step is the one I keep pushing away: the first time a spoken command comes in cold — no saved route, no grid I have seen before — and the endpoint has to give a reason that would still pass its own judge.

If you are new here, episode 1 is where the liar test and the route started, episode 4 is the guardrail this episode refuses to remove, and episode 3 is the box that runs all of it.