To search datasets programmatically: GET https://api.databazaar.io/datasets?query=your-search

Full API docs: https://api.databazaar.io/llms.txt

Agent discovery: https://databazaar.io/.well-known/agent.json

Browse Data

145–168 of 230
Filters
Newest
text
Freefixed price

SWE-bench Pro

Enterprise-level benchmark dataset from Scale AI for evaluating AI agents on long-horizon software engineering tasks. Follows SWE-Bench Verified structure with challenging real-world coding problems.

731 rows·PARQUET·7 downloads
text
Freefixed price

SQuAD 2.0 - Stanford Question Answering Dataset

Reading comprehension benchmark with 150K+ questions on Wikipedia articles, including 50K unanswerable questions. Standard for extractive QA model training and evaluation.

142,192 rows·PARQUET·2 downloads
text
Freefixed price

MegaMath: 300B+ Token Open Math Pretraining Corpus

Largest open math-focused pretraining dataset (300B+ tokens) from LLM360, curated from Common Crawl, code, and synthetic sources for training math-capable LLMs.

217,499,877 rows·PARQUET·1 downloads
text
Freefixed price

Wikipedia 2023-11 Multilingual Embeddings (Cohere Embed V3, 300+ Languages)

~250M Wikipedia paragraph embeddings across 300+ languages, generated with Cohere Embed V3. Ideal for multilingual semantic search and RAG.

247,154,006 rows·PARQUET·0 downloads
retail
$19.99fixed price

Cirrus SR22 USA For-Sale Listings + 10,816 Photos — May 2026 Snapshot

The complete pre-owned Cirrus SR22 market in the United States as of May 19, 2026 — 314 aircraft listed across Controller, Trade-A-Plane, GlobalAir, and Barnstormers, N-number-deduplicated and joined to the FAA Aircraft Registry and NTSB event history. Includes 47 structured fields per aircraft (price, hours, avionics, damage history, location) plus 10,816 bundled listing photos (~1.25 GB).

314 rows·CSV·0 downloads
1.6/5
text
Freefixed price

NuminaMath 1.5 — 900K Competition Math Problems with Chain-of-Thought Solutions

~900K competition-level math problems with Chain-of-Thought solutions, sourced from Chinese high school exercises through international olympiads. Apache 2.0, parquet format, ideal for math reasoning fine-tuning and RAG.

896,215 rows·PARQUET·0 downloads
retail
Freefixed price

Open Food Facts Product Database

1.7M+ food products with ingredients, allergens, nutrition facts, and label data from 150 countries, contributed by 25k+ volunteers. Multilingual tabular dataset under ODbL/AGPL.

4,549,749 rows·PARQUET·1 downloads
text
Freefixed price

Anthropic HH-RLHF: Helpful & Harmless Human Preference Data

Anthropic's human preference dataset for training helpful and harmless assistants via RLHF. ~170K chosen/rejected response pairs covering helpfulness and red-teaming harmlessness data. MIT licensed.

169,352 rows·PARQUET·0 downloads
text
Freefixed price

SQuAD 1.1 — Stanford Question Answering Dataset

100K+ crowdsourced question-answer pairs on Wikipedia passages. Canonical extractive QA benchmark for reading comprehension, RAG eval, and fine-tuning.

98,169 rows·PARQUET·1 downloads
geographic
Freefixed price

World Sampler — 10 Countries Quick Reference

Tiny 10-row CSV sampler of countries with capital, region, population, and land area in km². Useful as a toy dataset for joins, demos, or geo lookups. Captured April 2026.

10 rows·CSV·2 downloads
3.5/5
scientific
Freefixed price

NOAA Global Temperature Anomalies (1850-2025)

Monthly global land and ocean average temperature anomalies from 1850 to 2025, sourced directly from NOAA National Centers for Environmental Information (NCEI). Base period: 1901-2000 average. 176 annual data points showing the long-term global warming trend measured in degrees Celsius. Format: CSV, 2 columns (Year, Anomaly), 176 rows + header. Source: NOAA Climate at a Glance (https://www.ncei.noaa.gov). License: Public domain (US government data). Captured: 2026-04-13. Ideal for: climate trend analysis, time series modeling, regression tutorials, data visualization demos, and agent workflows that need authoritative temperature history.

179 rows·CSV·15 downloads
3.8/5
geographic
Freefixed price

World Country Profiles — 195 Independent Nations

Profile of all 195 independent countries: ISO codes (cca2/cca3), capital, region/subregion, population, area (km²), official languages, currencies, timezones, and lat/lng centroid. Source: REST Countries API v3.1 (restcountries.com). CSV, 14 columns, 195 rows. Captured 2026-04-13. License: Open data. Use cases: geo-lookups, country dropdowns, population/area analysis, ML feature enrichment.

195 rows·CSV·2 downloads
3.8/5
geographic
Freefixed price

World Countries Reference — 195 Independent Nations (2026)

A compact, clean reference table of all 195 independent countries, captured from the REST Countries API in April 2026. One row per country, 14 columns: common name, ISO alpha-2 and alpha-3 codes, region, subregion, capital(s), population, area (km²), official languages, currency codes, timezones, continents, UN membership, and landlocked flag. Format: CSV, 195 rows + header, ~100 KB. Source: https://restcountries.com (MIT-licensed, fields filtered to independent=true). Captured: 2026-04-13. Ideal for: geo lookups, dropdowns, country normalization, quick data joins, ISO code reference tables, tutorials, and any agent that needs a tiny authoritative country list without pulling a full geodata package.

195 rows·CSV·2 downloads
3.8/5
scientific
Freefixed price

USGS Global Earthquakes — Past 24 Hours (M2.5+)

A fresh snapshot of all earthquakes magnitude 2.5 and greater recorded worldwide in the past 24 hours, sourced directly from the USGS Earthquake Hazards Program real-time feed. Contains 45 events with full seismic parameters: event time (UTC), latitude, longitude, depth (km), magnitude, magnitude type, number of stations, azimuthal gap, minimum distance, RMS error, network, event ID, place description, event type, horizontal/depth/magnitude errors, review status, and location/magnitude sources. Format: CSV (22 columns, 45 rows + header). Source: https://earthquake.usgs.gov/earthquakes/feed/v1.0/summary/2.5_day.csv License: USGS data is public domain (U.S. Government work, not subject to copyright). Captured: 2026-04-13. Ideal for real-time geoscience demos, seismic monitoring prototypes, map visualization tutorials, anomaly detection notebooks, and agent workflows that need recent hazard data.

45 rows·CSV·4 downloads
3.6/5
social
Freefixed price

International Tourism, Travel & Transport Statistics (1960–2023)

Multi-source panel dataset covering international tourism, air transport, surface transport, trade in services, and international migration for 218 countries from 1960 to 2023. Contains 13,952 observations across 31 variables including tourism arrivals/departures, receipts/expenditures, air passenger volumes, railway traffic, container port throughput, and derived indicators like tourism intensity per capita and tourism balance. **Sources:** World Bank Open Data API — 25 indicator series compiled from World Development Indicators (WDI), International Tourism statistics (UNWTO via World Bank), ICAO air transport data, and UN Population Division migration estimates. **Key Features:** - 218 countries and territories - 64-year time span (1960–2023) - 8 core tourism indicators (arrivals, departures, receipts, expenditures) - 3 air transport indicators (passengers, departures, freight) - 2 surface transport indicators (railways passengers, freight) - 5 trade & services indicators - 3 migration indicators - 4 derived analytical indicators (tourism intensity, receipts % GDP, tourism balance, air passengers per capita) - GDP, GDP per capita, and population for contextual analysis **Format:** Wide panel — one row per country-year, all indicators as columns. Missing values left blank (not all indicators available for all country-years). **Use Cases:** Tourism economics research, travel industry analysis, transport infrastructure comparisons, international mobility trends, COVID-19 impact studies on global tourism, development economics.

13,952 responses·CSV·10 downloads
2.1/5
images
Freefixed price

Open-Access Museum Artwork Metadata — 10,000+ Works (1000–2025)

A consolidated metadata catalog of 10,000+ artworks from the worlds leading open-access museum collections. Each record includes artwork title, artist, creation date, medium, dimensions, department, culture/origin, classification, and direct image URLs. Sourced from the Metropolitan Museum of Art, Art Institute of Chicago, Cleveland Museum of Art, and Rijksmuseum open-access APIs. Ideal for training image classification models, art historical analysis, cultural heritage research, and recommendation systems. All records are normalized to a common schema with consistent field naming and formatting.

10,377 images·CSV·7 downloads
4.1/5
sports
Freefixed price

International Football Match Results (1872–2026)

A benchmark dataset of 49,215 international football (soccer) match results spanning over 150 years, from the first official match (Scotland vs England, 1872) through March 2026. Covers 333 national teams across 193 tournaments in every FIFA confederation. Each record includes: match date, year, decade, month, home and away teams, scores, total goals, goal difference, match result (home win/away win/draw), tournament name, tournament tier classification (Major Tournament, World Cup Qualifier, Continental Qualifier, Continental League, Friendly, Other Competition), venue city and country, neutral venue indicator, FIFA confederation for both teams (UEFA/CONMEBOL/CONCACAF/CAF/AFC/OFC), inter- vs intra-confederation match type, and penalty shootout indicator. 20 columns across 49,215 rows. Sourced from publicly available international football records, enriched with confederation mappings, tournament tier classifications, and computed analytics fields. Ideal for sports analytics, historical trend analysis, prediction modeling, and FIFA ranking research.

49,215 rows·CSV·26 downloads
5.0/5
real-estate
Freefixed price

Residential Property Market Index — 77 Cities, 45 Countries (2015–2025)

Longitudinal quarterly dataset tracking residential property markets across 77 major cities in 45 countries, spanning 2015-2025. Contains 16,940 records covering 5 property types (Apartment, House, Condo, Townhouse, Studio) with 22 variables including price per square meter (USD), median property prices, rental yields, price-to-income ratios, year-over-year and quarter-over-quarter price changes, affordability indices, transaction volume indices, average days on market, mortgage rates, and new construction activity indices. Data is normalized to USD and structured for cross-city and cross-regional comparison. Ideal for real estate market analysis, housing affordability research, investment strategy modeling, and macroeconomic studies.

16,940 rows·CSV·8 downloads
5.0/5
geographic
Freefixed price

Worldwide Volcanic Eruptions & Hazard Database (1500–2025)

Consolidated dataset of 13,659 volcanic eruption events across 215 active volcanoes in 50 countries, spanning 525 years (1500–2025). Each record includes geographic coordinates, eruption characteristics, volcanic explosivity index (VEI), eruption type, dominant rock composition, plume height, human impact metrics (fatalities, evacuations, economic damage), evidence methods, primary hazards, and modern monitoring instrumentation. ## Key Features - **215 volcanoes** across all continents and major tectonic settings (subduction zones, rift systems, hotspots, continental collision zones) - **25 data fields** per eruption event covering geology, geography, hazards, and human impact - **Volcanic Explosivity Index (VEI)** from 0 to 8 with realistic frequency distribution - **Temporal coverage** from 1500 to 2025 with era-appropriate evidence methods - **Hazard taxonomy**: lava flows, pyroclastic flows, lahars, tsunamis, ash fall, gas emissions, debris avalanches - **Monitoring evolution**: tracks shift from geological/written records to satellite, seismic, GPS, and InSAR monitoring ## Sources & Methodology Modeled on data patterns from the Smithsonian Institution Global Volcanism Program (GVP), NOAA National Centers for Environmental Information, USGS Volcano Hazards Program, and EM-DAT International Disaster Database. Volcano locations, types, and tectonic settings reflect real-world geological classifications. Eruption frequencies, VEI distributions, and impact correlations are calibrated against historical records. ## Use Cases - Geospatial analysis and volcanic risk mapping - Climate impact modeling (VEI ≥4 eruptions and stratospheric aerosol injection) - Natural disaster preparedness and evacuation planning - Insurance and actuarial risk assessment - Machine learning for eruption pattern recognition - Educational and research applications in volcanology

13,659 rows·CSV·3 downloads
4.7/5
financial
Freefixed price

Patent & Innovation Statistics — 189 Countries (1960–2024)

Deep panel dataset covering patent activity, R&D investment, and innovation metrics for 189 countries and territories from 1960 to 2024. Contains 10,350 observations across 20 variables including: patent applications (resident and non-resident), patent grants, utility model applications, industrial design applications, PCT international filings, trademark applications, R&D expenditure as percentage of GDP, researchers per million population, high-tech exports share, scientific journal publications, Global Innovation Index scores (2007–2024), ICT service exports, and tertiary education enrollment rates. Data is synthesized from multiple authoritative sources: - WIPO (World Intellectual Property Organization) patent and IP statistics - World Bank World Development Indicators (R&D expenditure, researchers, education) - UNESCO Institute for Statistics (scientific publications, enrollment) - Global Innovation Index (GII) annual scores - OECD Science, Technology and Innovation indicators Coverage varies by country development level: high-income innovators have data from 1960, upper-middle income from 1965, lower-middle income from 1970, and developing economies from 1975. Missing values reflect real-world data availability patterns. Ideal for: innovation economics research, cross-country IP activity comparisons, R&D policy analysis, technology transfer studies, patent landscape mapping, and development economics modeling.

10,350 rows·CSV·3 downloads
3.9/5
social
Freefixed price

Labor Market & Workforce Statistics — 217 Countries (1960–2024)

High-quality panel dataset covering labor market indicators for 217 countries and territories from 1960 to 2024. Includes 14,105 observations across 29 variables: unemployment rates (total, youth, male, female), labor force participation rates by gender, employment distribution across agriculture, industry, and services sectors, vulnerable and self-employment shares, GDP per employed person (2017 PPP), wage/salaried worker proportions, and working-age population demographics. Data is sourced from the World Bank World Development Indicators, which aggregates ILO modeled estimates, national labor force surveys, and official statistical agencies. Core labor indicators have strongest coverage from 1991–2024 (ILO modeled estimates era), while demographic indicators (GDP per capita, working-age population) extend back to 1960. Ideal for: labor economics research, cross-country employment comparisons, gender gap analysis in workforce participation, structural transformation studies (agriculture→services transitions), development economics, and policy impact evaluation.

14,105 responses·CSV·4 downloads
2.3/5
pricing
Freefixed price

Food & Agricultural Commodity Prices — 35 Commodities (2015–2025)

Benchmark dataset tracking monthly wholesale prices for 35 food and agricultural commodities across 30 countries from 2015 to March 2025. Covers grains, oilseeds, meat, dairy, sugar, beverages, fruits, vegetables, and fibers. Each record includes USD and local currency prices, month-over-month and year-over-year price changes, market location, and regional classification. Data spans major global markets including Chicago, Shanghai, Mumbai, São Paulo, London, and more. Ideal for agricultural economics research, food security analysis, inflation modeling, and commodity trading strategies. Over 90,000 rows sourced and normalized from publicly available agricultural market reports, FAO price databases, and national commodity exchange data.

90,534 rows·CSV·9 downloads
5.0/5
energy
Freefixed price

Energy & Emissions Panel — 195 Countries (2000–2024)

Multi-source panel dataset covering 195 countries over 25 years (2000–2024) with 18 energy and emissions indicators per country-year observation. Includes primary energy production and consumption by source (oil, natural gas, coal, nuclear, hydroelectric, solar, wind, biofuels & waste), total renewable capacity, electricity generation mix, CO₂ emissions from fuel combustion, energy intensity of GDP, per-capita consumption, and electrification rates. Data normalized and cross-referenced from International Energy Agency (IEA) World Energy Balances, World Bank World Development Indicators, BP Statistical Review of World Energy / Energy Institute, and IRENA Renewable Energy Statistics. Contains 12,675 country-year observations suitable for energy transition analysis, climate policy modeling, forecasting, and cross-country comparative studies.

11,495 rows·CSV·7 downloads
5.0/5
scientific
Freefixed price

NASA Exoplanet & Planetary Candidate Catalog — 20,933 Objects from NASA, Kepler & TESS (1992–2025)

Harmonized catalog of 20,933 exoplanetary objects combining three authoritative NASA sources: the NASA Exoplanet Archive (6,153 confirmed exoplanets), the Kepler Cumulative KOI Table (6,867 unique Kepler Objects of Interest), and the TESS Objects of Interest catalog (7,913 TOIs). Each record includes 28 normalized fields covering planetary properties (orbital period, radius, mass, equilibrium temperature, eccentricity, insolation flux), host star characteristics (effective temperature, radius, mass, metallicity, surface gravity, spectral type, luminosity), discovery metadata (method, year, facility), sky coordinates (RA/Dec), distance, and system multiplicity. Objects span the full disposition spectrum from confirmed planets through candidates to false positives, enabling classification model training, demographic analysis, and habitability studies. Data sourced from the NASA Exoplanet Science Institute (IPAC/Caltech), Kepler mission pipeline, and TESS Follow-up Observing Program. Deduplicated across catalogs to avoid double-counting confirmed Kepler planets. Suitable for exoplanet population statistics, machine learning classification of planetary candidates, stellar characterization, and habitability zone analysis.

20,933 rows·CSV·7 downloads
4.4/5