building

GSM Symbolic Contamination

Is the GSM-Symbolic math benchmark now memorized by post-2024 models?

PythonLLM Evaluation

An evaluation-integrity study testing whether post-2024 language models have memorized the GSM-Symbolic math benchmark, using symbolic variants to detect contamination.

What I built