Live Osprey scorecard

Which models should the agent fleet use next?

Not a generic leaderboard. This radar re-ranks BenchLM evidence for Osprey work: multi-agent ops, websites/dashboards, Google/Meta ads, SEO research, and video production support. Subscription-first scoring: cost is nearly ignored. OpenAI Codex, Grok, and Kimi are first-class for Hermes/MoA. Claude/Anthropic is Claude Code only — never in Mixture of Agents.

Host: models.osprey.solutions Subscription-first Role fleets Claude = Code only MoA ready
Loading

Fetching weighted model feed…

Equal usage lanes

Hermes agents, OpenClaw agents, and Claude Code are equal. Claude never joins Hermes/OpenClaw MoA.

Default Hermes / OpenClaw top 3

MoA-safe shortlist only. Claude Code has its own equal lane above.

Role fleets

Different jobs get different model mixes. Ads and video are not the same problem as general Hermes ops.

Default Mixture of Agents

Generalist MoA from the default top 3. Role fleets override this when the job is specialized.

Tracked models

Slim table of what matters. Hide noise. Sort by Osprey score.

# Model Osprey Lane Agentic Coding Multi Know Ops $ in/out Context
Loading…