How well do agents handle fixed-income tasks?
FIBRE scores frontier models on 202 practitioner-written fixed-income tasks across two tracks — answering from a written brief, and researching filings live with tools — and reports what each answer costs in seconds and dollars.
Leaderboard
Score weights every task equally regardless of track, so a model scored on more of them carries more of its own result. The # column reports standing on whatever the board is ranked on — that score, or the category chosen above it — and keeps reporting it whichever column you sort by. Question-answering and Agentic-research break that score down by track and are scored against the same rubric per task. Latency is the median wall clock per task and cost is the mean at list API prices from the v1-2026-08 table; a model with no published price shows no cost rather than a guess.
Ranked by lab, the board keeps one row per lab: that lab’s best model on Score, with every figure still that one model’s. Best rather than an average, which would only penalise a lab for also entering its small and cheap models.
Trade-offs
Best per task is the diamond: for each of the 200 tasks every model attempted, the highest score any of them got, averaged — 95.0 against 91.3 for opus-5 used on everything. Nobody can build it. The choice is made with the result in hand, which a router does not have when it has to choose, so it marks the edge of what having 5 different models is worth rather than a score anyone can reach. It is also flattered by noise — the highest of 5 measurements rises with the count even when the models behind them do not improve — so it will drift up as models are added and is not comparable between boards with different numbers of them. 2 tasks are left out of it, for not having been attempted by every model. Its cost is measured the way the board measures a model’s, on the question-answering track that holds more of the tasks.