Lab app · The work
Everything the lab is doing.
Experiments and support work are different classes and stay that way — one has a hypothesis and a kill condition, the other has a counterpart team and a deliverable. Neither is what anybody actually does on a Tuesday. Underneath both sits the same thing: a harness somebody has to build, an eval somebody has to write, a dataset somebody has to get cleared, and a write-up somebody has to sign.
Every run by stage with its owner, pair, days in stage and overruns visibly ageing, and beside it the other half — work done with somebody else, which has a counterpart and a deliverable instead of a hypothesis. This is what the lab is doing, and it opens here because a board is readable in a second and a grouped list is not.
Proposed
hypothesis, kill condition, cost
Voting
Type 3 and 4 only · scores locked before discussion
Does it still agree with itself?
Temporal decision memory for a claims agent, measured on the same case two weeks apart
Learning AgentsRunning
own harness, findings logged from the session
The skill wiki
Compiling what an agent learns across sessions into a persistent, compiled skill store
Personal WikiMeasuring
evals, measurement runs
Concluded
one of five terminal states
Memory bake-off
Four memory designs on a 40-session support corpus, scored on decision consistency rather than recall
Agentic Memory SystemWhere the tokens go
A token cost ledger across six client patterns, attributing every token to a cache state
Cost redux on tokensThe judge on trial
LLM-as-judge scored against a human panel, because agreement is checked and not assumed
Eval HarnessesHow fast is fast enough
The latency floor for a voice agent on AU telco infrastructure, and what it costs to reach it
Voice and VisionThe other half
Work done with somebody else.
Standing up the gateway, model access with procurement, making AI-SDLC repeatable. No hypothesis and no kill condition — a counterpart team and a deliverable instead, so it gets its own columns rather than being bent into a run. Handed over is the success state: work the lab still owns a year later has failed at the thing it was for.
8 experiments in flight against 5 open contributions, 84 days committed. The charter protects experiment time structurally; a floor of 50% is where that stops being a slogan.
Stalled · 1
Model and tool approval with Procurement. Waiting on lifecycle accountability, which the plan dates to September and nobody has yet claimed. The lab can supply the evaluation method and cannot decide who signs. Carrying it as active would have been dishonest and would have hidden the ask.
Spotted, not yet agreed to. Costs nothing until somebody commits.
- MLwith Platform Engineering
The lab has agreed to be the dependency, with an owner and an effort against it.
- MLwith Partnerships
In flight, drawing on the same attention the shortlist does.
- DYwith Platform Engineering, Security30d quiet
- AWwith Partnerships, Procurement23d quiet
- MMwith Engineering, Internal transformation squad16d quiet
Identified costs nothing until somebody commits, which is the point of the column: the plan names the lab as a dependency on more initiatives than it can carry, and taking them silently is how a lab ends up running four things it never agreed to. Attribution to an objective is a commit with a name on it, because it is how budget gets justified.
2026-09-03