Model Router
One library every AI project calls: picks the model and effort per task, logs every dollar, proves it with evals.
building
Summary
A thin Python wrapper around the Anthropic SDK that resolves a named route (such as the morning brief or a financial extraction step) to a model, effort level, caching policy, and fallback. Every call lands in a SQLite ledger with usage and cost. An eval harness sweeps model and effort per route and picks the cheapest configuration that still passes.
Problem
Every project I built hardcoded one model for every call and never logged what it cost. I had no idea whether the morning brief cost five cents or five dollars, and no way to know if a cheaper setting would have been fine.
AI Usage
The LLM does the task; the router does everything around it deterministically: routing policy from a config file, cost math from a pricing table, validators that decide whether to escalate, and an eval sweep that measures pass rate against cost per completed task. Routing decisions are made by measured evals, not by another model call guessing difficulty.
- Built: the morning brief agents already run through it, logging model, tokens, and cost for every call.
- Built: named routes, cache-health reporting, and a policy gate checked before any network call. 49 tests passing.
- Planned: effort and model sweep per route with a Pareto chart of score vs cost.
- Planned: escalation measured honestly against a single strong model at low effort, and dropped if it loses.
Stack
Python, Anthropic SDK, SQLite, YAML