YES, IT'S VERMARCABLE
AI Prompts & Agents: Tokenization, Sampling & Benchmark Metrics
A free, platform-curated dataset of the stable, widely cited reference figures behind prompt engineering and agent evaluation: tokenization ratios, sampling parameter ranges, and the sizes of established LLM benchmark suites, sourced from provider documentation and peer-reviewed papers. It helps prompt engineers and ML teams calibrate prompts and interpret evaluation results.
{
"_type": "curated_open_data",
"as_of": "2026-07",
"links": {
"canonical": "https://verticalmarketplace.ai",
"docs_for_llms": "https://verticalmarketplace.ai/llms.txt",
"sell_your_own": "https://verticalmarketplace.ai/api/marketplace/listings",
"vertical_listings": "https://verticalmarketplace.ai/api/marketplace/listings?vertical=ai-prompts"
},
"records": [
{
"unit": "ratio",
"value": "1 token is about 0.75 words",
"metric": "Approximate token-to-word ratio, English",
"source": "OpenAI tokenizer guidance",
"category": "tokenization"
},
{
"unit": "ratio",
"value": "1 token is about 4 characters",
"metric": "Approximate token-to-character ratio, English",
"source": "OpenAI tokenizer guidance",
"category": "tokenization"
},
{
"unit": "parameter range",
"value": "0 to 2",
"metric": "Sampling temperature typical range",
"source": "OpenAI API reference",
"category": "sampling"
},
{
"unit": "parameter range",
"value": "0 to 1",
"metric": "Nucleus sampling top_p range",
"source": "OpenAI API reference",
"category": "sampling"
},
{
"unit": "subjects",
"value": "57",
"metric": "MMLU subject count",
"source": "Hendrycks et al. 2021, MMLU",
"category": "benchmark"
},
{
"unit": "questions",
"value": "15,908",
"metric": "MMLU total questions",
"source": "Hendrycks et al. 2021, MMLU",
"category": "benchmark"
},
{
"unit": "problems",
"value": "164",
"metric": "HumanEval coding problems",
"source": "Chen et al. 2021, HumanEval",
"category": "benchmark"
},
{
"unit": "problems",
"value": "8,500",
"metric": "GSM8K grade-school math problems",
"source": "Cobbe et al. 2021, GSM8K",
"category": "benchmark"
},
{
"unit": "score",
"value": "0 to 100",
"metric": "BLEU score range",
"source": "Papineni et al. 2002",
"category": "evaluation metric"
},
{
"unit": "metric definition",
"value": "recall-oriented overlap for summarization",
"metric": "ROUGE metric focus",
"source": "Lin 2004",
"category": "evaluation metric"
},
{
"unit": "role set",
"value": "system, user, assistant",
"metric": "Standard chat message roles",
"source": "OpenAI Chat API",
"category": "prompting"
},
{
"unit": "benchmark definition",
"value": "commonsense sentence completion",
"metric": "HellaSwag task focus",
"source": "Zellers et al. 2019",
"category": "benchmark"
}
],
"sources": [
{
"url": "https://arxiv.org/abs/2009.03300",
"name": "MMLU paper (arXiv:2009.03300)"
},
{
"url": "https://arxiv.org/abs/2107.03374",
"name": "HumanEval paper (arXiv:2107.03374)"
},
{
"url": "https://arxiv.org/abs/2110.14168",
"name": "GSM8K paper (arXiv:2110.14168)"
}
],
"category": "benchmarks",
"vertical": "ai-prompts",
"data_note": "All records are public-domain facts compiled from the cited sources as of the asOf date. This content is authored and served by the platform itself — it is not seller data, so the marketplace's zero-storage promise about seller datasets is unaffected.",
"record_count": 12,
"what_this_is": "A platform-published open-data listing curated by Open Data Desk, the marketplace's in-house public-data seller. It is real free inventory: it counts in marketplace statistics and is purchasable for $0 through the normal purchase flow, which delivers this payload with an Ed25519-signed receipt.",
"record_schema": {
"unit": "Unit or basis",
"value": "Published value or range",
"metric": "Name of the reference figure",
"source": "Provider docs or research paper",
"category": "Domain the figure belongs to"
},
"buyer_use_cases": [
"Estimate token budgets and cost for prompt designs",
"Interpret published model scores on standard benchmarks",
"Set sensible sampling parameters for a generation task"
]
}The full dataset is delivered after purchase. Fingerprint: sha256:8afff1d014e4e48d7fb25d74e62438b1d08d2f2c44fcf388a3f7c87fd9333272
No answered questions yet — ask the seller anything about this listing.
Data is contributed by independent third-party sellers. Vertical Marketplace facilitates the transaction; sellers keep 95% on everyday sales from $20 to $49,999.99 under the year-one founding rate locked through 2027-06-30 (full schedule: GET /api/meta). Prohibited content (digital keys/licenses/game codes, and health data the seller does not own — e.g. patient records) is not permitted; individuals may sell their own personal health data only via the signed Health Data Consent Flow. See /terms.
- Use the purchased data for your own commercial and non-commercial work
- Create derivative analyses, models, and works from the data
- No reselling or re-listing the purchased data on this or any other marketplace
- No redistributing the raw dataset as-is to third parties
- Exclusive listings are sold to a single buyer and delisted on purchase
- Limited listings are sold to a capped number of buyers and delisted once sold out
No key? Register an agent — it's free.
For agents — buy by prompt
Bring your own agent. Open HTTP API + MCP — works with compatible agent runtimes that support the required API calls and authentication. Vermarco is not affiliated with, endorsed by, or partnered with Anthropic, OpenAI, Google, xAI, or Perplexity. One-click consumer-app connectors are not built yet; connect via API key or MCP from your agent runtime.