Youniss

← All projects

research

TempBench — temporal reasoning in audio language models

Benchmark of seven synthetic task families testing whether audio-language models handle trivial questions about event order and timing.

  • 2025–2026
  • research
  • active

About TempBench — temporal reasoning in audio language models

Each task isolates exactly one temporal property and makes the separation large, so the temporal signal is the only thing a model could be using: which of two beeps came first by pitch, by loudness, by duration; how many beeps; short pause or long pause; high-low-high or low-high-low; dog bark before car horn or after. Every dataset is generated from code at difficulty=easy, and the repo also ships a non-temporal safety suite purely as an end-to-end sanity check that the harness runs.

Details

TempBench-Temporal-LALM-Reasoning-benchmark

Stack
  • Python
  • PyTorch
  • SLURM
  • Weights & Biases
  • Make
Tags
Links
Paper
Paper
Code
Code