Live Osprey scorecard

Which models should the agent fleet use next?

Not a generic leaderboard. This radar re-ranks multi-source evidence for Osprey work: multi-agent ops, websites/dashboards, Google/Meta ads, SEO research, and video production support. Subscription-first scoring: cost is nearly ignored. OpenAI Codex, Grok, Kimi, and Claude (Anthropic) are first-class under the right runtime. Hermes primary MoA stays non-Claude, with Claude allowed as a reference/reviewer. OpenClaw can select Claude via Claude OAuth CLI. Claude Code remains the equal dedicated Claude lane. Scores blend five sources: BenchLM (35%) · LMArena (30%) · Artificial Analysis (20%) · LiveBench (8%) · llm-stats (7%).

Host: models.osprey.solutions 5-source blend Subscription-first OpenClaw Claude OAuth Hermes Claude reference MoA ready
Loading

Fetching weighted model feed…

Equal usage lanes

Hermes primary MoA (non-Claude) · OpenClaw may select Claude via OAuth CLI · Claude Code equal lane · Hermes Claude reference

Default Hermes / OpenClaw top 3

Hermes primary MoA shortlist (non-Claude). OpenClaw and Claude Code have their own lanes above.

Role fleets

Different jobs get different model mixes. Ads and video are not the same problem as general Hermes ops.

Default Mixture of Agents

Generalist MoA from the default top 3. Role fleets override this when the job is specialized.

Tracked models

Slim table of what matters. Hide noise. Sort by Osprey score.

# Model Osprey BenchLM Arena AA LiveB Stats Lane Agentic Coding Multi Know Context
Loading…

Sources & methodology

Five independent quality sources are blended by their configured weights. OpenRouter supplies availability and context metadata only.

Loading source methodology…

How the Osprey score works

The score is a confidence-adjusted, renormalized blend of whichever scoring sources have evidence for a model. Missing sources are shown as “—”; they are not silently treated as zero. OpenRouter has 0% score weight and only confirms catalog availability and context metadata.