Datasets
Open-source datasets for ML research and development, published on Hugging Face.
- Triton Kernel Code DatasetTriton kernel bodies and CUDA→Triton translation pairs for fine-tuning code models.
- Aesthetic Images7.9k+ filtered high-aesthetic images for diffusion and image-text alignment models.
- Brazilian Portuguese TTS≈150 hours of multi-speaker Brazilian Portuguese for text-to-speech synthesis.
- Obama Voice Sample Dataset25+ minutes of clean 24kHz speech from public addresses, optimized for RVC training.
- Speech Dataset InternationalMultilingual speech dataset for TTS and speech synthesis research.