
TAT-QA: Hybrid Table-Text QA Benchmark for Financial Annual Report Reasoning
TAT-QA's hybrid table-text questions showed evidence grounding, not arithmetic, is finance AI's bottleneck. Fine-tuned 7B LLMs hit 83% F1 by 2024.

TAT-QA's hybrid table-text questions showed evidence grounding, not arithmetic, is finance AI's bottleneck. Fine-tuned 7B LLMs hit 83% F1 by 2024.

FinQA found neural models scored 61% on financial-report math versus 91% for human experts, collapsing to 22% on three-or-more-step programs.

FinanceBench tests 16 AI setups on 10,231 real SEC filing questions: shared-vector-store RAG answers only 19% right, so retrieval is not the bottleneck.

DSPy's compiler lifted Llama2-13b from 9.4% to 46.9% on GSM8K, pointing finance AI pipelines toward maintainable declarative LLM calls.

LATS unifies ReAct, Tree of Thoughts, and Reflexion in one MCTS framework, hitting 92.7% pass@1 on HumanEval with GPT-4.

Self-RAG trains an LLM to decide when to retrieve and self-grade results, hitting 55.8% on PopQA and 80.2 FactScore — beating ChatGPT on five benchmarks.

Voyager's persistent code skill library discovers 3.3× more Minecraft items than prior SOTA without fine-tuning—the reuse pattern ledger agents need.

HippoRAG's OpenIE+PageRank memory hits 89.1% Recall@5 on 2WikiMultiHopQA versus 68.2% ColBERTv2—path-aware retrieval for multi-year ledgers.

AgentBench scored GPT-4 4.01 versus 0.96 for the best open-source LLM. Those failures are exactly what breaks a Beancount write-back agent on a live ledger.

BloombergGPT trained on 569B financial tokens, yet GPT-4 matched it with no finance pretraining. For accounting agents, tool-use beats model internals.

AutoGen's two-agent conversation lifts MATH accuracy from 55% to 69%, and its SafeGuard agent adds up to 35 F1 points on unsafe-code detection.

Gorilla's Retriever-Aware Training cuts LLM API hallucination rates from 78% to 11%, making tool calls reliable enough for finance agents that write entries.