Benchmark methodology
Provider-reported, third-party, and FrontierDot-measured scores are stored and displayed separately. They are never merged into one unlabeled leaderboard.
A FrontierDot run records exact model version, prompt/dataset version, sampling parameters, runner region, latency, tokens, estimated cost, harness commit, artifact hash, limitations, and status. Historical runs are append-only.
Evidence labels
provider_reported— published by the model provider, with a source URL.third_party_reported— published by an external evaluator, with a source URL.measured_by_frontierdot— run by the FrontierDot harness, with methodology and reproducibility metadata.
Initial families (v0.1)
FrontierDot Coding 0.1.0
Versioned repository tasks that require reading a small codebase, applying a multi-file change, and repairing tests. Scoring is task-level pass/fail plus instruction-following checks. Historical runs are append-only.
Dimensions: repo_understanding, bug_fixing, multi_file_change, test_repair, instruction_following
- v0.1 is a small private task set, not a public league table
- Results are only published with exact model version, harness commit, and runner region
- English-first task statements in v0.1
FrontierDot Tool Use 0.1.0
Scripted tool-calling scenarios with a fixed tool schema. The harness scores correct tool choice, argument validity, recovery after a tool error, and valid structured output. Provider-reported tool scores are stored separately.
Dimensions: tool_selection, argument_construction, sequential_tool_use, failure_recovery, structured_outputs
- Does not evaluate arbitrary plugin ecosystems
- Tool set is FrontierDot-owned and versioned with the harness
FrontierDot Multilingual 0.1.0
Locale-native writing, extraction, business communication, and localization tasks. This is not a translation contest. English leaderboards are not used as a proxy for other locales.
Dimensions: native_writing, extraction, business_communication, localization, terminology, instruction_understanding
- v0.1 covers English, Chinese, Japanese, and Korean
- Human review is required before a measured score is published
FrontierDot Real-World Inference 0.1.0
Practical inference tasks that resemble production LLM use: summarization, extraction, classification, transformation, planning, and long-context synthesis. Cost and latency are recorded with the score.
Dimensions: summarization, document_extraction, classification, data_transformation, seo_content_workflows, agent_planning, long_context_synthesis
- Workload mix is FrontierDot-defined and will change with harness versions
- Latency is runner-region specific and is not a global ranking
Update policy
Retesting appends a new run. Published historical measurements are not overwritten. Scores without exact model version, date, and evidence type are not shown as FrontierDot measurements.
Cost and latency are runner-region specific. They are not a global ranking.