Not a generic leaderboard. This radar re-ranks multi-source evidence for Osprey work: multi-agent ops, websites/dashboards, Google/Meta ads, SEO research, and video production support. Subscription-first scoring: cost is nearly ignored. OpenAI Codex, Grok, Kimi, and Claude (Anthropic) are first-class under the right runtime. Hermes primary MoA stays non-Claude, with Claude allowed as a reference/reviewer. OpenClaw can select Claude via Claude OAuth CLI. Claude Code remains the equal dedicated Claude lane. Scores blend five sources: BenchLM (35%) · LMArena (30%) · Artificial Analysis (20%) · LiveBench (8%) · llm-stats (7%).
Fetching weighted model feed…
Hermes primary MoA (non-Claude) · OpenClaw may select Claude via OAuth CLI · Claude Code equal lane · Hermes Claude reference
Hermes primary MoA shortlist (non-Claude). OpenClaw and Claude Code have their own lanes above.
Different jobs get different model mixes. Ads and video are not the same problem as general Hermes ops.
Generalist MoA from the default top 3. Role fleets override this when the job is specialized.
Slim table of what matters. Hide noise. Sort by Osprey score.
| # | Model | Osprey | BenchLM | Arena | AA | LiveB | Stats | Lane | Agentic | Coding | Multi | Know | Context |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Loading… | |||||||||||||
Five independent quality sources are blended by their configured weights. OpenRouter supplies availability and context metadata only.
The score is a confidence-adjusted, renormalized blend of whichever scoring sources have evidence for a model. Missing sources are shown as “—”; they are not silently treated as zero. OpenRouter has 0% score weight and only confirms catalog availability and context metadata.