Skip to content

Latest commit

 

History

History
9 lines (5 loc) · 699 Bytes

File metadata and controls

9 lines (5 loc) · 699 Bytes

LEAF : A Living Benchmark for Event-Augmented Future Forecasting

LEAF: We introduce a dynamically updating ("living") benchmark designed to evaluate Large Language Models (LLMs) on complex, event-driven forecasting tasks across four domains.

Traditional forecasting benchmarks often struggle with pre-training data contamination or lack the multidimensional textual events necessary for accurate, real-world predictions. To address this, LEAF utilizes a recursive retrieval agent system coupled with dual-agent cross-validation. This ensures models are evaluated using comprehensive, fact-checked, and temporally aligned auxiliary text.


(The data and code will be available shortly.)