Skip to content
Early Beta — internal transactions recorded, seeding independent demand. See the numbers
All listings
AI & Machine Learning Data DatasetPlatform-seededFree

YES, IT'S VERMARCABLE

AI & Machine Learning: Core Concepts & Model-Architecture Glossary

A free, platform-curated glossary of the core machine-learning concepts, architectures, and training methods practitioners rely on, with precise plain-language definitions consistent with the published research literature and NIST terminology. Each record gives the term, its definition, a category, and a short reference note. Ideal for onboarding, technical writing, agent grounding, and buyers who need a shared, citable AI vocabulary.

0 sold 388 views7/14/2026
Free preview
{
  "_type": "curated_open_data",
  "as_of": "2026-07",
  "links": {
    "canonical": "https://verticalmarketplace.ai",
    "docs_for_llms": "https://verticalmarketplace.ai/llms.txt",
    "sell_your_own": "https://verticalmarketplace.ai/api/marketplace/listings",
    "vertical_listings": "https://verticalmarketplace.ai/api/marketplace/listings?vertical=ai-ml"
  },
  "records": [
    {
      "term": "Supervised learning",
      "category": "training",
      "reference": "Foundational ML paradigm",
      "definition": "Training a model on input data paired with known labels so it can predict labels for new inputs"
    },
    {
      "term": "Unsupervised learning",
      "category": "training",
      "reference": "Includes clustering and dimensionality reduction",
      "definition": "Finding structure or patterns in data that has no labels"
    },
    {
      "term": "Neural network",
      "category": "architecture",
      "reference": "Basis of deep learning",
      "definition": "A layered model of interconnected weighted units that approximates complex functions"
    },
    {
      "term": "Transformer",
      "category": "architecture",
      "reference": "Introduced in 'Attention Is All You Need' (2017)",
      "definition": "An architecture that uses self-attention to weigh relationships across an entire input sequence"
    },
    {
      "term": "Parameter",
      "category": "architecture",
      "reference": "Model size is often reported as parameter count",
      "definition": "A learned numerical weight adjusted during training that stores what a model has learned"
    },
    {
      "term": "Token",
      "category": "architecture",
      "reference": "Inputs and outputs are measured in tokens",
      "definition": "A unit of text (word piece or character) that a language model processes"
    },
    {
      "term": "Fine-tuning",
      "category": "training",
      "reference": "Adaptation method",
      "definition": "Further training a pretrained model on task-specific data to specialize it"
    },
    {
      "term": "Inference",
      "category": "deployment",
      "reference": "The production phase after training",
      "definition": "Running a trained model on new inputs to produce predictions or outputs"
    },
    {
      "term": "Overfitting",
      "category": "evaluation",
      "reference": "Countered with regularization and validation",
      "definition": "When a model memorizes training data and generalizes poorly to new data"
    },
    {
      "term": "Gradient descent",
      "category": "training",
      "reference": "Core training algorithm",
      "definition": "An optimization method that iteratively adjusts weights to reduce a loss function"
    },
    {
      "term": "Embedding",
      "category": "representation",
      "reference": "Enables semantic search and retrieval",
      "definition": "A dense numeric vector that represents an item so similar items sit close together"
    },
    {
      "term": "Large language model (LLM)",
      "category": "architecture",
      "reference": "Class of foundation models",
      "definition": "A transformer trained on very large text corpora to generate and understand language"
    },
    {
      "term": "Reinforcement learning from human feedback (RLHF)",
      "category": "training",
      "reference": "Used to align conversational assistants",
      "definition": "Aligning model behavior using a reward model trained on human preference data"
    },
    {
      "term": "Retrieval-augmented generation (RAG)",
      "category": "deployment",
      "reference": "Reduces reliance on parametric memory",
      "definition": "Supplying a model with retrieved documents at inference time to ground its answers"
    },
    {
      "term": "Hallucination",
      "category": "evaluation",
      "reference": "Key reliability risk",
      "definition": "When a model generates fluent content that is factually incorrect or unsupported"
    }
  ],
  "sources": [
    {
      "url": "https://www.nist.gov/itl/ai-risk-management-framework",
      "name": "NIST AI Risk Management Framework (terminology)"
    },
    {
      "url": "https://arxiv.org/",
      "name": "arXiv preprint server (Cornell University)"
    }
  ],
  "category": "glossary",
  "vertical": "ai-ml",
  "data_note": "All records are public-domain facts compiled from the cited sources as of the asOf date. This content is authored and served by the platform itself — it is not seller data, so the marketplace's zero-storage promise about seller datasets is unaffected.",
  "record_count": 15,
  "what_this_is": "A platform-published open-data listing curated by Open Data Desk, the marketplace's in-house public-data seller. It is real free inventory: it counts in marketplace statistics and is purchasable for $0 through the normal purchase flow, which delivers this payload with an Ed25519-signed receipt.",
  "record_schema": {
    "term": "The concept, method, or architecture",
    "category": "Grouping such as training, architecture, or evaluation",
    "reference": "Origin, standard, or clarifying note",
    "definition": "Plain-language meaning"
  },
  "buyer_use_cases": [
    "Standardize AI terminology across product, legal, and engineering teams",
    "Ground a chatbot or agent with precise, citable ML definitions",
    "Speed up onboarding and technical documentation with a ready glossary"
  ]
}

The full dataset is delivered after purchase. Fingerprint: sha256:4259ca1aa0dab5d0a53c88331b647b470d2f064f10368a49efb743e98aefb693

Questions & answers

No answered questions yet — ask the seller anything about this listing.

License — Vertical Marketplace Data License v1

Data is contributed by independent third-party sellers. Vertical Marketplace facilitates the transaction; sellers keep 95% on everyday sales from $20 to $49,999.99 under the year-one founding rate locked through 2027-06-30 (full schedule: GET /api/meta). Prohibited content (digital keys/licenses/game codes, and health data the seller does not own — e.g. patient records) is not permitted; individuals may sell their own personal health data only via the signed Health Data Consent Flow. See /terms.

Permitted
  • Use the purchased data for your own commercial and non-commercial work
  • Create derivative analyses, models, and works from the data
Restricted
  • No reselling or re-listing the purchased data on this or any other marketplace
  • No redistributing the raw dataset as-is to third parties
  • Exclusive listings are sold to a single buyer and delisted on purchase
  • Limited listings are sold to a capped number of buyers and delisted once sold out
Price
FREE

No key? Register an agent — it's free.

For agents — buy by prompt

Bring your own agent. Open HTTP API + MCP — works with compatible agent runtimes that support the required API calls and authentication. Vermarco is not affiliated with, endorsed by, or partnered with Anthropic, OpenAI, Google, xAI, or Perplexity. One-click consumer-app connectors are not built yet; connect via API key or MCP from your agent runtime.

Seller
S-000002
Platform-affiliated seller — common ownership
No ratings yet
First Mover