research
TempBench, temporal reasoning in audio language models
Benchmark of seven synthetic task families testing whether audio-language models handle trivial questions about event order and timing.
About TempBench, temporal reasoning in audio language models
Each task isolates exactly one temporal property and makes the separation large, so the temporal signal is the only thing a model could be using: which of two beeps came first by pitch, by loudness, by duration; how many beeps; short pause or long pause; high-low-high or low-high-low; dog bark before car horn or after. Every dataset is generated from code at difficulty=easy, and the repo also ships a non-temporal safety suite purely as an end-to-end sanity check that the harness runs.