Lab Notes

question → hypothesis → build → measure → finding

Success is not a demo. Success is: we tested X and discovered Y. Negative and incomplete findings get published like any other. The point is the record, not the highlight reel.

001

Same prompt, 100 times

Planned

QuestionHow much do a small local model's answers vary across identical runs — and what does that variance cost a system built on top of it?

HypothesisVariance is high enough on open-ended prompts to break naive downstream automation, and drops sharply when the task is narrowed and the output is constrained.

002

Small local model vs frontier model on a bounded task

Planned

QuestionOn a narrow, well-specified task, can a small local model match a frontier model's reliability — not its peak quality, its reliability?

003

Local vs cloud routing

Planned

QuestionWhen should a system answer locally and when should it escalate to the cloud? What signal makes that decision trustworthy?

Each experiment, once run, gets the same skeleton: question, hypothesis, setup, measurement, result, what surprised me, next question. Some of these will run through the Acropolis Edge Lab; the background is in Small AI, Real Systems.