Labs
Small things you can actually use.
A finding you can only read is a claim. A finding you can run is an argument. Each of these is a working version of something the lab tested — play with it, then look at the eval suite underneath it and decide for yourself whether we have earned the recommendation.
- Toys
- 0
- Eval cases
- 0
- Failing
- 0
The failing cases are shown, not hidden. A suite with nothing red in it is either trivial or not being read.
Also shipped
Recommendations that came with something you can run.
Not every finding gets a toy. These are the tested recommendations that shipped with a repo, an eval or a demo attached — the same records as the library, filtered to the ones with an artifact.
Use a structured episodic store with summarised recall, not a raw vector memorymemory-adapter — one store, three harnessesTestedRoute through a gateway you control; do not standardise on a vendor's gardenRouting table walkthrough: three gardens, six task typesTestedEvery LLM judge ships with a human agreement score or does not shipJudge calibration panel v2 — 300 items, three task typesTestedPut backpressure on agent fan-out before you put it on the modelbounded-fanout — orchestrator queue wrapperTested