Skip to content
FrontierDot

Benchmark methodology

Provider-reported, third-party, and FrontierDot-measured scores are stored and displayed separately. They are never merged into one unlabeled leaderboard.

A FrontierDot run records exact model version, prompt/dataset version, sampling parameters, runner region, latency, tokens, estimated cost, harness commit, artifact hash, limitations, and status. Historical runs are append-only.

Evidence labels

Initial families (v0.1)

FrontierDot Coding 0.1.0

Versioned repository tasks that require reading a small codebase, applying a multi-file change, and repairing tests. Scoring is task-level pass/fail plus instruction-following checks. Historical runs are append-only.

Dimensions: repo_understanding, bug_fixing, multi_file_change, test_repair, instruction_following

FrontierDot Tool Use 0.1.0

Scripted tool-calling scenarios with a fixed tool schema. The harness scores correct tool choice, argument validity, recovery after a tool error, and valid structured output. Provider-reported tool scores are stored separately.

Dimensions: tool_selection, argument_construction, sequential_tool_use, failure_recovery, structured_outputs

FrontierDot Multilingual 0.1.0

Locale-native writing, extraction, business communication, and localization tasks. This is not a translation contest. English leaderboards are not used as a proxy for other locales.

Dimensions: native_writing, extraction, business_communication, localization, terminology, instruction_understanding

FrontierDot Real-World Inference 0.1.0

Practical inference tasks that resemble production LLM use: summarization, extraction, classification, transformation, planning, and long-context synthesis. Cost and latency are recorded with the score.

Dimensions: summarization, document_extraction, classification, data_transformation, seo_content_workflows, agent_planning, long_context_synthesis

Update policy

Retesting appends a new run. Published historical measurements are not overwritten. Scores without exact model version, date, and evidence type are not shown as FrontierDot measurements.

Cost and latency are runner-region specific. They are not a global ranking.