research
TempBench — temporal reasoning in audio language models
Benchmark of seven synthetic task families testing whether audio-language models handle trivial questions about event order and timing.
About TempBench — temporal reasoning in audio language models
Each task isolates exactly one temporal property and makes the separation large, so the temporal signal is the only thing a model could be using: which of two beeps came first by pitch, by loudness, by duration; how many beeps; short pause or long pause; high-low-high or low-high-low; dog bark before car horn or after. Every dataset is generated from code at difficulty=easy, and the repo also ships a non-temporal safety suite purely as an end-to-end sanity check that the harness runs.
Details
- Stack
- Python
- PyTorch
- SLURM
- Weights & Biases
- Make
- Tags
- Links