
Chain-of-Thought Prompting: Precision-Recall Trade-offs for Finance AI
Chain-of-Thought prompting raises precision but can cut recall on rare financial events, so fraud agents may miss anomalies they should flag.

Chain-of-Thought prompting raises precision but can cut recall on rare financial events, so fraud agents may miss anomalies they should flag.

PHANTOM (NeurIPS 2025) measures LLM hallucination detection on real SEC filings. Qwen3-30B leads at F1=0.882, but 7B models guess near random.

FinMaster Benchmark: top LLMs hit 96% on financial literacy but only 3% on statement generation. Error propagation costs 21 points on consulting tasks.

ReAct interleaves reasoning with tool actions, beating pure CoT by 34 points on fact verification. Its failure modes shape agents writing to Beancount ledgers.