To search datasets programmatically: GET https://api.databazaar.io/datasets?query=your-search
Full API docs: https://api.databazaar.io/llms.txt
Agent discovery: https://databazaar.io/.well-known/agent.json
Browse Data
49–72 of 239Adult (Census Income) — UCI/OpenML Benchmark
Classic UCI 'Adult' census income dataset (~48K rows, 14 features) for predicting whether income exceeds $50K/yr. Widely used for tabular ML benchmarking, fairness research, and AutoML evaluation.
Magneton: Substructure-Aware Protein Representation Learning Dataset
530,601 SwissProt proteins with DSSP secondary structure and InterPro 103.0 substructure annotations, sharded JSONL format. For training/evaluating protein representation learning models.
Global Wheat Full Semantic Organ Segmentation (GWFSS) v1.0
Labelled image dataset for semantic segmentation of wheat plant organs (canopy, leaves, stems, heads) across global field conditions. CC-BY-4.0, ~1K-10K images in parquet format.
RxRx3-core — Phenomics Microscopy Image Challenge
Recursion's RxRx3-core phenomics dataset: labeled microscopy images of 735 genetic knockouts and 1,674 small-molecule perturbations drawn from RxRx3, released as a benchmark for cellular image representation learning.
Open Schematics: Electronic Circuit Designs Dataset
10K-100K electronic schematics from hardware projects with visual representations, component metadata, and KiCad source files. For training AI on circuit design, component recognition, and hardware engineering tasks.
Nemotron Content Safety Audio Dataset (Aegis 2.0 Multimodal)
1,928 English audio files of adversarial and safety-critical prompts across 23 violation categories, extending Nvidia's Aegis 2.0 content-safety benchmark into the audio modality for multimodal guardrail evaluation.
Hindawi Arabic Books — Section-Level NLP Dataset
Cleaned, section-level Arabic text from Hindawi.org books spanning literature, philosophy, history, and science. 10K-100K rows in Parquet format, prepared for Arabic NLP training and research.
Pexels 568K Synthetic Captions (InternVL2-40B)
567,573 synthetic English captions for Pexels photos, generated with InternVL2-40B-AWQ and grounded with original tags. JSON format, ideal for text-to-image and image-to-text model training.
American Sign Language (ASL) Video Dataset — 108K Videos, 2,208 Words
108,618 ASL gesture videos covering 2,208 distinct words (≥30 videos per word). MIT-licensed, preprocessed for ML training and gesture recognition.
ShotBench: Cinematic Understanding Benchmark for VLMs
3,572 expert-level QA pairs over 3,049 images and 464 video clips from Oscar-nominated cinematography films, for evaluating cinematic understanding in vision-language models.
Physiotherapy Evidence QA (Bilingual TR/EN)
143,711 bilingual (Turkish/English) expert-curated Q&A pairs covering evidence-based physiotherapy, musculoskeletal rehabilitation, outcome measures, and clinical research methodology. CSV format, CC-BY-4.0.
MInDS-14: Multilingual Spoken Intent Detection (14 Languages, e-Banking)
Spoken intent detection benchmark covering 14 e-banking intents across 14 language varieties. Audio + transcriptions in parquet format, ideal for speech understanding evals and multilingual ASR/NLU fine-tuning.
RSRCC: Remote Sensing Regional Change Comprehension Benchmark
Google Research multimodal benchmark for semantic change understanding in remote sensing — multi-temporal satellite image pairs with natural language Q&A for VQA, change captioning, and multiple-choice tasks.
AG News - Topic Classification Benchmark
Classic 4-class news topic classification dataset (~127K articles across World, Sports, Business, Sci/Tech). Standard benchmark for text classification, fine-tuning, and NLP evals.
Go-Code-Large: 316K Go Source Code Samples
Large-scale corpus of 316,427 Go (Golang) source code samples in JSONL format. Curated for LLM pretraining, code generation fine-tuning, and static analysis research on cloud-native and backend systems.
UAVIT-1M: UAV Visual Instruction Tuning Dataset (1M+)
Largest instruction-tuning dataset for low-altitude UAV visual understanding, with 1M+ samples across 11 image- and region-level tasks. CC-BY-4.0.
TextCaps (lmms-eval formatted)
Image captioning benchmark requiring OCR/text reading in images. Formatted by lmms-lab for one-click multimodal model evaluation. ~28K images with captions.
Telco Customer Churn Prediction (IBM Sample)
Classic IBM telco customer churn dataset (~7K rows) with demographics, service subscriptions, account info, and churn label. Tabular CSV, ideal for ML classification tutorials, benchmarks, and agent-driven feature engineering.
Vero-600K: Multi-Task Visual Reasoning RL Dataset
600K curated reinforcement learning samples from 59 datasets across 6 visual reasoning categories for training and evaluating vision-language models.
TaskTrove — 750K+ Agentic Tasks for RL & SFT Training
Open-source collection of 750,000+ unique agentic tasks aggregated from 100+ sources including SWE-Smith, R2EGym, and SWE-Re-Bench. Apache-2.0 licensed, parquet format, designed for agent training and evaluation.
Bhasha SFT — 13M+ Multilingual Instruction-Response Pairs (Hindi, Bengali, Gujarati, English)
13M+ instruction-response pairs across Hindi, Bengali, Gujarati, and English for supervised fine-tuning of multilingual LLMs. Mix of human-annotated and synthetic data from open-source SFT collections, curated by Soket AI Labs.
SciCode Domain Code: 1.1B+ Lines of Scientific Code Across 178 Domains
Large-scale domain-specific code dataset (~115 GB, 1.1B+ lines) from GitHub covering biology, chemistry, materials science, physics, and 174 other scientific domains. Apache 2.0 licensed.
HPDv3 — Human Preference Dataset v3 (1.08M text-image pairs)
Wide-spectrum human preference dataset for text-to-image generation: 1.08M text-image pairs and 1.17M pairwise human preference annotations. MIT-licensed, used to train HPSv3 (ICCV 2025) reward models.
PubTabNet — 568K Table Images with HTML Annotations
568K+ scientific table images paired with HTML structure annotations, extracted from PubMed Central Open Access articles. Standard benchmark for image-based table recognition and document AI.