cavendish

Share graph · provenance running forward

ML

Michelle Lam

Analytics Lead · evals · lab

What they dropped, which clusters it landed in, what got elected, and which experiments and findings descended from it. Everyone’s is visible to everyone, symmetrically. Show outcomes, not counts — no totals, no rankings, no rollups.

Dropped

23

signals with this person attached

Landed in

16

distinct clusters

Elected

11

of those, now fields

Descended

24

experiments and published items

Drops

What descended

recommendationTested

Use a structured episodic store with summarised recall, not a raw vector memory

recommendationTested

Route through a gateway you control; do not standardise on a vendor's garden

recommendationTested

Prompt caching: use for stable prefixes over 2k tokens; expect 30–45%, not 60%

recommendationTested

Every LLM judge ships with a human agreement score or does not ship

recommendationTested

Open-weight models for classification and extraction; frontier for agentic loops

recommendationTested

Distil to a small model only after the frontier baseline is measured on the same eval

standing answerTested

Which model for structured extraction?

standing answerTested

What does inference actually cost right now?

standing answerTested

Retrieval or fine-tuning for this?

standing answerTested

Which memory layer should a new agent use?

standing answerAssessed

What eval tooling do we use?

positionAssessed

Sovereign inference and the end of US default

positionAssessed

Learning without weights: where continual learning actually lands

experiment · concludedValidated

Memory bake-off

experiment · concludedValidated

Right model, right task

experiment · concludedValidated

Where the tokens go

experiment · concludedRefuted

The judge on trial

experiment · concludedSuperseded

Close enough to switch?

experiment · measuring

The canary suite

experiment · measuring

Small model, back of the store

experiment · measuring

Let the agent break it

experiment · running

The skill wiki

experiment · voting

Does it still agree with itself?

experiment · proposed

Learning without retraining

Follow this person’s finds

Following someone whose drops are consistently good is the internal version of the external voice watchlist, and often a better source than any detector.