FrontierFinance: A Benchmark for Financial Reasoning Agents
FrontierFinance is an open benchmark measuring how well AI systems support the full investment workflow — from screening and discovery to earnings analysis and portfolio monitoring. Queries are open-ended financial research tasks scored against expert-written rubric items using a rubric-based evaluation methodology.
Methodology: Each query in FrontierFinance represents a realistic financial research task. Model responses are evaluated against a set of expert-defined rubric criteria, each marked as must-have or informational. The benchmark rewards completeness and precision rather than surface-level keyword matching.
Dataset:220 public queries spanning use cases including financial data & modeling, sector, industry & macro analysis, earnings & events, company research, coverage & catalyst monitoring, and screening & discovery. The dataset is available on Hugging Face and the evaluation code is open-source on GitHub.
Citation
@article{zhang2026frontierfinance,
title = {{FrontierFinance}: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents},
author = {Yuhao Zhang and O. Ozan Koyluoglu and Thejas Venkatesh and Richard Diehl Martinez and Vishank Bhatia and Arash Alidoust and Ashwin Paranjape},
journal = {arXiv preprint arXiv:2608.11683},
year = {2026}
}