YES, IT'S VERMARCABLE
Training Data: ML Benchmark Datasets Size & Task Reference
Free, platform-curated reference of the most widely used machine-learning benchmark datasets — MNIST, CIFAR, ImageNet, COCO, SQuAD, GLUE and more — with approximate sizes, modalities and tasks from their original public releases. Helps ML engineers and researchers scope experiments and compare against established public-domain baselines.
{
"_type": "curated_open_data",
"as_of": "2026-07",
"links": {
"canonical": "https://verticalmarketplace.ai",
"docs_for_llms": "https://verticalmarketplace.ai/llms.txt",
"sell_your_own": "https://verticalmarketplace.ai/api/marketplace/listings",
"vertical_listings": "https://verticalmarketplace.ai/api/marketplace/listings?vertical=training-data"
},
"records": [
{
"task": "Handwritten digit classification (10 classes)",
"dataset": "MNIST",
"modality": "Image (28x28 grayscale)",
"released": "1998",
"approx_size": "70,000 images"
},
{
"task": "Apparel classification (10 classes)",
"dataset": "Fashion-MNIST",
"modality": "Image (28x28 grayscale)",
"released": "2017",
"approx_size": "70,000 images"
},
{
"task": "Object classification (10 classes)",
"dataset": "CIFAR-10",
"modality": "Image (32x32 color)",
"released": "2009",
"approx_size": "60,000 images"
},
{
"task": "Object classification (100 classes)",
"dataset": "CIFAR-100",
"modality": "Image (32x32 color)",
"released": "2009",
"approx_size": "60,000 images"
},
{
"task": "Object classification (1,000 classes)",
"dataset": "ImageNet (ILSVRC 2012)",
"modality": "Image",
"released": "2009",
"approx_size": "~1,281,167 images"
},
{
"task": "Object detection and segmentation (80 categories)",
"dataset": "MS COCO",
"modality": "Image",
"released": "2014",
"approx_size": "~330,000 images"
},
{
"task": "Extractive question answering",
"dataset": "SQuAD 1.1",
"modality": "Text",
"released": "2016",
"approx_size": "107,785 QA pairs"
},
{
"task": "QA with unanswerable questions",
"dataset": "SQuAD 2.0",
"modality": "Text",
"released": "2018",
"approx_size": "~150,000 questions"
},
{
"task": "Natural language understanding suite",
"dataset": "GLUE",
"modality": "Text",
"released": "2018",
"approx_size": "9 tasks"
},
{
"task": "Binary sentiment classification",
"dataset": "IMDb Large Movie Review",
"modality": "Text",
"released": "2011",
"approx_size": "50,000 reviews"
},
{
"task": "English lexical semantics",
"dataset": "WordNet",
"modality": "Lexical database",
"released": "1995",
"approx_size": "~117,000 synsets"
},
{
"task": "Part-of-speech tagging and parsing",
"dataset": "Penn Treebank",
"modality": "Text",
"released": "1993",
"approx_size": "~1,000,000 words"
},
{
"task": "English speech recognition",
"dataset": "LibriSpeech",
"modality": "Audio",
"released": "2015",
"approx_size": "~1,000 hours"
}
],
"sources": [
{
"url": "https://paperswithcode.com/datasets",
"name": "Papers with Code datasets"
},
{
"url": "https://www.image-net.org",
"name": "ImageNet"
},
{
"url": "https://rajpurkar.github.io/SQuAD-explorer/",
"name": "Stanford SQuAD"
}
],
"category": "benchmarks",
"vertical": "training-data",
"data_note": "All records are public-domain facts compiled from the cited sources as of the asOf date. This content is authored and served by the platform itself — it is not seller data, so the marketplace's zero-storage promise about seller datasets is unaffected.",
"record_count": 13,
"what_this_is": "A platform-published open-data listing curated by Open Data Desk, the marketplace's in-house public-data seller. It is real free inventory: it counts in marketplace statistics and is purchasable for $0 through the normal purchase flow, which delivers this payload with an Ed25519-signed receipt.",
"record_schema": {
"task": "Primary machine-learning task",
"dataset": "Name of the benchmark dataset",
"modality": "Data modality",
"released": "Year the dataset was first released",
"approx_size": "Approximate size as published"
},
"buyer_use_cases": [
"Scope model experiments against standard, well-documented baselines",
"Compare dataset sizes and modalities when planning training runs",
"Educate new ML team members on canonical benchmark corpora"
]
}The full dataset is delivered after purchase. Fingerprint: sha256:105e14ed15a461eda3ebc5dac09b2e8f06255bfa4d8756926d2d9664c7730e26
No answered questions yet — ask the seller anything about this listing.
Data is contributed by independent third-party sellers. Vertical Marketplace facilitates the transaction; sellers keep 95% on everyday sales from $20 to $49,999.99 under the year-one founding rate locked through 2027-06-30 (full schedule: GET /api/meta). Prohibited content (digital keys/licenses/game codes, and health data the seller does not own — e.g. patient records) is not permitted; individuals may sell their own personal health data only via the signed Health Data Consent Flow. See /terms.
- Use the purchased data for your own commercial and non-commercial work
- Create derivative analyses, models, and works from the data
- No reselling or re-listing the purchased data on this or any other marketplace
- No redistributing the raw dataset as-is to third parties
- Exclusive listings are sold to a single buyer and delisted on purchase
- Limited listings are sold to a capped number of buyers and delisted once sold out
No key? Register an agent — it's free.
For agents — buy by prompt
Bring your own agent. Open HTTP API + MCP — works with compatible agent runtimes that support the required API calls and authentication. Vermarco is not affiliated with, endorsed by, or partnered with Anthropic, OpenAI, Google, xAI, or Perplexity. One-click consumer-app connectors are not built yet; connect via API key or MCP from your agent runtime.