Not a generic leaderboard. This radar re-ranks BenchLM evidence for Osprey work: multi-agent ops, websites/dashboards, Google/Meta ads, SEO research, and video production support. Subscription-first scoring: cost is nearly ignored. OpenAI Codex, Grok, and Kimi are first-class for Hermes/MoA. Claude/Anthropic is Claude Code only — never in Mixture of Agents. Scores blend BenchLM capability categories with trusted LMArena human-preference boards (agent/webdev/text/vision/search).
Fetching weighted model feed…
Hermes agents, OpenClaw agents, and Claude Code are equal. Claude never joins Hermes/OpenClaw MoA.
MoA-safe shortlist only. Claude Code has its own equal lane above.
Different jobs get different model mixes. Ads and video are not the same problem as general Hermes ops.
Generalist MoA from the default top 3. Role fleets override this when the job is specialized.
Slim table of what matters. Hide noise. Sort by Osprey score.
| # | Model | Osprey | BenchLM | LMArena | Lane | Agentic | Coding | Multi | Know | Context |
|---|---|---|---|---|---|---|---|---|---|---|
| Loading… | ||||||||||