ai-ml
Embed2Image contrastive retrieval
Audio-text retrieval experiments on Clotho comparing a PaSST+RoBERTa baseline against a pseudo-image ViT encoder head.
About Embed2Image contrastive retrieval
As part of a course, I participated in the DCASE 2025 Challenge with a small team. We tried many different approaches and were able to work on the Vienna Scientific Computing cluster. While we didn’t win, we got a lot of experience with audio ML, and our technical report was published on the DCASE website.
Details
- Stack
- Python
- PyTorch Lightning
- PaSST
- RoBERTa
- Weights & Biases
- Tags
- Links
More in ai-ml
See all- nanoBeard
A small pirate-themed GPT trained from scratch on a piratized TinyStories corpus and then SFT-tuned, with a versioned training codebase.
- PaperNavigator
Academic paper discovery tool that profiles a research question, snowballs through citations, filters with an LLM and generates a report.
- DISCO
CLIP-based classifier for detecting implicit suggestive imagery, trained on a hand-labelled fashion dataset and published on Hugging Face.
- Falcon-Twig
Fine-tuning run for Falcon H1 7B on tool calling that underperformed the base instruct model, written up in a post.