Youniss

← All projects

ai-ml

Embed2Image contrastive retrieval

Audio-text retrieval experiments on Clotho comparing a PaSST+RoBERTa baseline against a pseudo-image ViT encoder head.

  • 2025
  • ai-ml
  • experiment

About Embed2Image contrastive retrieval

As part of a course, I participated in the DCASE 2025 Challenge with a small team. We tried many different approaches and were able to work on the Vienna Scientific Computing cluster. While we didn’t win, we got a lot of experience with audio ML, and our technical report was published on the DCASE website.

Details

embed2image-contrastive-retrieval

Stack
  • Python
  • PyTorch Lightning
  • PaSST
  • RoBERTa
  • Weights & Biases
Tags
Links
Paper
Paper
Datasets
audio-text-embed-to-images2,360 downloads
Code
Code

More in ai-ml

See all
  • nanoBeard
    • 2026
    • shipped

    A small pirate-themed GPT trained from scratch on a piratized TinyStories corpus and then SFT-tuned, with a versioned training codebase.

  • PaperNavigator
    • 2026
    • shipped

    Academic paper discovery tool that profiles a research question, snowballs through citations, filters with an LLM and generates a report.

  • DISCO
    • 2025
    • shipped

    CLIP-based classifier for detecting implicit suggestive imagery, trained on a hand-labelled fashion dataset and published on Hugging Face.

  • Falcon-Twig
    • 2025
    • shipped

    Fine-tuning run for Falcon H1 7B on tool calling that underperformed the base instruct model, written up in a post.