Share graph · provenance running forward
Michelle Lam
Analytics Lead · evals · lab
What they dropped, which clusters it landed in, what got elected, and which experiments and findings descended from it. Everyone’s is visible to everyone, symmetrically. Show outcomes, not counts — no totals, no rankings, no rollups.
Dropped
signals with this person attached
Landed in
distinct clusters
Elected
of those, now fields
Descended
experiments and published items
Drops
What descended
Use a structured episodic store with summarised recall, not a raw vector memory
Route through a gateway you control; do not standardise on a vendor's garden
Prompt caching: use for stable prefixes over 2k tokens; expect 30–45%, not 60%
Every LLM judge ships with a human agreement score or does not ship
Open-weight models for classification and extraction; frontier for agentic loops
Distil to a small model only after the frontier baseline is measured on the same eval
Which model for structured extraction?
What does inference actually cost right now?
Retrieval or fine-tuning for this?
Which memory layer should a new agent use?
What eval tooling do we use?
Sovereign inference and the end of US default
Learning without weights: where continual learning actually lands
Memory bake-off
Right model, right task
Where the tokens go
The judge on trial
Close enough to switch?
The canary suite
Small model, back of the store
Let the agent break it
The skill wiki
Does it still agree with itself?
Learning without retraining
Follow this person’s finds
Following someone whose drops are consistently good is the internal version of the external voice watchlist, and often a better source than any detector.