
Chain-of-Thought Prompting: Precision-Recall Trade-offs for Finance AI
Chain-of-Thought prompting raises precision but can cut recall on rare financial events, so fraud agents may miss anomalies they should flag.
#llm
Large language model research with applications in financial tasks

Chain-of-Thought prompting raises precision but can cut recall on rare financial events, so fraud agents may miss anomalies they should flag.

PHANTOM (NeurIPS 2025) measures LLM hallucination detection on real SEC filings. Qwen3-30B leads at F1=0.882, but 7B models guess near random.

FinMaster Benchmark: top LLMs hit 96% on financial literacy but only 3% on statement generation. Error propagation costs 21 points on consulting tasks.

ReAct interleaves reasoning with tool actions, beating pure CoT by 34 points on fact verification. Its failure modes shape agents writing to Beancount ledgers.

Toolformer teaches a 6.7B model to call APIs via perplexity filtering, beating GPT-3 175B on arithmetic. Its single-step design blocks chained ledger calls.

FinBen finds GPT-4 at 0.63 exact match on FinQA and 0.54 on stock forecasting, barely above random, so accounting agents need validation, not raw LLM math.