← all repositories
Chengjun023/agent-smith

Stop paying GPT-5 rates to fix commas

Agent Smith routes routine Codex work to cheaper models and shows the receipts — a local router plus a native macOS usage monitor.

agent-smith
Collecting fresh signals — velocity needs a few days of history.
collecting data…
star history

What it does

Two pieces in one repo: an adaptive router that picks a Codex model per task based on a local difficulty heuristic (1.0–10.0, in 0.1 steps), and Codex Float, a native macOS floating window that tracks tasks, subagents, token usage, and reference costs. The router sends easy work to lighter models and reserves the expensive ones for hard cases. Everything runs locally — scoring, metering, and price comparisons add no model calls.

The interesting bit

The author is unusually honest about epistemics: missing usage or prices stay “unknown” rather than zero, prices are frozen at snapshot time instead of retroactively repriced, and acceptance/rejection carry explicit evidence semantics. The pilot benchmark is deliberately tiny — six synthetic tasks × three strategies — precisely so you can cross-examine it.

Key highlights

  • Router pilot: 86.89% less reference cost vs fixed Astra, with all 18/18 tasks passing deterministic graders
  • Codex Float deduplicates usage across restarts and log rotation via its own SQLite ledger; conversation bodies stay out
  • Model discovery from models.dev on a TTL, with last-good snapshots — discovery never grants execution access
  • Python parts use only the standard library; benchmarks can run offline via demo and replay modes
  • 130 router tests and 50 Float tests in the imported baseline

Caveats

  • Execution currently only works through Codex; other providers are catalog metadata, not usable backends
  • The pilot is one attempt on six synthetic tasks — the README itself says broader quality equivalence needs separate experiments
  • Difficulty scoring is an uncalibrated heuristic, and the fixed-6.1-Sol comparison had 24,832 cached input tokens the others didn’t, which skews cost math

Verdict

Worth a look if you use Codex on macOS and want cost visibility with honest accounting. If you’re hoping for multi-provider routing or calibrated difficulty scoring, that case hasn’t been tried yet — and the README admits it.

Frequently asked

What is Chengjun023/agent-smith?
Agent Smith routes routine Codex work to cheaper models and shows the receipts — a local router plus a native macOS usage monitor.
Is agent-smith open source?
Yes — Chengjun023/agent-smith is open source, released under the MIT license.
What language is agent-smith written in?
Chengjun023/agent-smith is primarily written in Python.
How popular is agent-smith?
Chengjun023/agent-smith has 508 stars on GitHub.
Where can I find agent-smith?
Chengjun023/agent-smith is on GitHub at https://github.com/Chengjun023/agent-smith.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.