Key takeaways

Public benchmarks are marketing-adjacent signals, not purchase orders. They rarely capture your tool schemas, bilingual product tone, P95 under load, or token bills. Use benches to build a shortlist of 2–3 China LLMs, then run private evals. Swift Horse does not publish a live “best model” rank.

What benches miss

Tool-calling parity, streaming + tools, rate limits from overseas, CN/EN mixed prompts, RAG citation faithfulness, and cost per successful task. Agent costs explode beyond chat benches—see /en/articles/china-llm-agent-tool-calling-2026 and pricing.

Private eval recipe

Freeze 20–50 tasks from production → identical system prompts → log pass/fail, human preference, tokens, latency → include 5 tool tasks → estimate monthly cost → pick primary + failover (/en/articles/china-llm-latency-failover-2026).

Next steps on Swift Horse

Shortlist /en/articles/best-chinese-llm-2026 → matrix /en/articles/top-chinese-ai-models-2026 → services /en/services → models /en/models.

FAQ

Which China LLM ranks #1 on public benches?

It rotates by suite and date. Treat #1 as a shortlist hint, not a contract.

Can I skip private evals?

Not for production. Benches miss your tools, latency, and cost shape.

How do benchmarks relate to API pricing?

High scores with expensive output tokens can lose on TCO—see /en/articles/china-llm-api-pricing-2026.

Does Swift Horse host a leaderboard?

No live ranked leaderboard—structured model pages and selection guides only.