cern-opendata-mcp-server

v0.1.1 pre-1.0

Search CERN Open Data, fetch records, files, analysis environments, CMS good-run lists, HLT paths via MCP. STDIO or Streamable HTTP.

cern-opendata.caseyjhand.com/mcp
claude mcp add --transport http cern-opendata-mcp-server https://cern-opendata.caseyjhand.com/mcp
codex mcp add cern-opendata-mcp-server --url https://cern-opendata.caseyjhand.com/mcp
{
  "mcpServers": {
    "cern-opendata-mcp-server": {
      "url": "https://cern-opendata.caseyjhand.com/mcp"
    }
  }
}
gemini mcp add --transport http cern-opendata-mcp-server https://cern-opendata.caseyjhand.com/mcp
{
  "mcpServers": {
    "cern-opendata-mcp-server": {
      "command": "bunx",
      "args": [
        "mcp-remote",
        "https://cern-opendata.caseyjhand.com/mcp"
      ]
    }
  }
}
{
  "mcpServers": {
    "cern-opendata-mcp-server": {
      "type": "http",
      "url": "https://cern-opendata.caseyjhand.com/mcp"
    }
  }
}
curl -X POST https://cern-opendata.caseyjhand.com/mcp \
  -H "Content-Type: application/json" \
  -H "Accept: application/json, text/event-stream" \
  -d '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-11-25","capabilities":{},"clientInfo":{"name":"curl","version":"1.0.0"}}}'

Tools

7

cern_opendata_search_records

open-world

Search the CERN Open Data Portal's datasets, software, environments, documentation and supplementary records with exact-vocabulary filters and an optional full-text query. The filters are experiment, record type, collision energy and type, file format, data-taking year, event count, availability and collection. Returns compact hits with recids plus live facet counts. Each facet ignores its own filter, so its counts show the alternatives under the other filters. Filter values are exact upstream; common spellings are normalized, and cern_opendata_list_reference lists the vocabulary. Paging reaches the first 10,000 matches.

read
invocation
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "cern_opendata_search_records",
    "arguments": {}
  }
}
schema
{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "type": "object",
  "properties": {
    "query": {
      "description": "Full-text query, sent verbatim as an OpenSearch query_string (AND between terms; title matches weigh double). Field forms such as title:\"…\", recid:(1 OR 2) or run_period:(\"Run2012B\") work; cern_opendata_list_reference with topic query_syntax lists them. Omit to browse by filters alone.",
      "type": "string",
      "maxLength": 500
    },
    "type": {
      "description": "Record types (OR): a primary (Dataset, Documentation, Environment, Software, Supplementaries, News) or Primary::Secondary such as Dataset::Collision or Dataset::Simulated. Array or comma-separated string, up to 7. Omit for every served type; Glossary is not served.",
      "maxItems": 7,
      "type": "array",
      "items": {
        "type": "string",
        "maxLength": 100,
        "description": "One type value."
      }
    },
    "experiment": {
      "description": "Experiments (OR): ALICE, ATLAS, CMS, DELPHI, JADE, LHCb, OPERA, PHENIX, TOTEM. Array or comma-separated string, up to 9; case is normalized.",
      "maxItems": 9,
      "type": "array",
      "items": {
        "type": "string",
        "maxLength": 100,
        "description": "One experiment value."
      }
    },
    "collision_energy": {
      "description": "Collision energies (OR), such as 7TeV, 8TeV, 13TeV or 5.02TeV. Array or comma-separated string, up to 15; \"13TeV, 13.6TeV\" is one upstream value and is kept whole.",
      "maxItems": 15,
      "type": "array",
      "items": {
        "type": "string",
        "maxLength": 100,
        "description": "One collision_energy value."
      }
    },
    "collision_type": {
      "description": "Collision types (OR): pp, PbPb, pPb, e+e-, Interfill. PbPb also matches the Pb-Pb spelling. Array or comma-separated string, up to 6.",
      "maxItems": 6,
      "type": "array",
      "items": {
        "type": "string",
        "maxLength": 100,
        "description": "One collision_type value."
      }
    },
    "file_type": {
      "description": "File formats and data tiers (OR), such as nanoaod, miniaod, aod, root, DAOD_PHYSLITE or csv. Array or comma-separated string, up to 20; case is normalized for known values.",
      "maxItems": 20,
      "type": "array",
      "items": {
        "type": "string",
        "maxLength": 100,
        "description": "One file_type value."
      }
    },
    "year_from": {
      "description": "Earliest data-taking year, inclusive. Alone, it means this year onward; for one year, set year_from and year_to to it.",
      "type": "integer",
      "minimum": 1900,
      "maximum": 2100
    },
    "year_to": {
      "description": "Latest data-taking year, inclusive. Alone, it means up to this year.",
      "type": "integer",
      "minimum": 1900,
      "maximum": 2100
    },
    "min_events": {
      "description": "Minimum number of events, inclusive.",
      "type": "integer",
      "minimum": 0,
      "maximum": 9007199254740991
    },
    "max_events": {
      "description": "Maximum number of events, inclusive.",
      "type": "integer",
      "minimum": 0,
      "maximum": 9007199254740991
    },
    "availability": {
      "description": "Record availability (OR): online, partial, ondemand (on tape, requested before download) or requested. Array or comma-separated string, up to 4.",
      "maxItems": 4,
      "type": "array",
      "items": {
        "type": "string",
        "maxLength": 100,
        "description": "One availability value."
      }
    },
    "collection": {
      "description": "Portal collections (OR), exact and case-sensitive, such as CMS-Validated-Runs; copy spellings from a record's collections field. Array or comma-separated string, up to 10.",
      "maxItems": 10,
      "type": "array",
      "items": {
        "type": "string",
        "maxLength": 100,
        "description": "One collection name, exact and case-sensitive."
      }
    },
    "sort": {
      "description": "bestmatch (relevance), mostrecent (newest first), title (A-Z) or title_desc (Z-A). Omit for the portal default: bestmatch with a query, mostrecent without.",
      "type": "string",
      "enum": [
        "bestmatch",
        "mostrecent",
        "title",
        "title_desc"
      ]
    },
    "limit": {
      "description": "Hits per page, 1-50.",
      "default": 10,
      "type": "integer",
      "minimum": 1,
      "maximum": 50
    },
    "page": {
      "description": "Page number, from 1. page × limit may not exceed 10,000.",
      "default": 1,
      "type": "integer",
      "minimum": 1,
      "maximum": 9007199254740991
    }
  },
  "required": [
    "limit",
    "page"
  ],
  "additionalProperties": false
}
view source ↗

cern_opendata_get_records

open-world

Fetch full metadata for 1-20 records in one call, by recid, DOI, CMS dataset path (/Primary/Era/TIER) or documentation slug. Returns the description, run periods, collision and distribution details, related records, a software-environment summary, the license and a ready citation. Documentation and news pages include their markdown body, cut at 30,000 characters. File lists are not included; use cern_opendata_list_files. Identifiers that do not resolve come back under missing with guidance; they do not fail the call.

read
invocation
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "cern_opendata_get_records",
    "arguments": {
      "ids": "<ids>"
    }
  }
}
schema
{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "type": "object",
  "properties": {
    "ids": {
      "description": "Identifiers to resolve, 1-20: an array, or one comma-separated string. Forms may be mixed; duplicates collapse.",
      "minItems": 1,
      "maxItems": 20,
      "type": "array",
      "items": {
        "type": "string",
        "maxLength": 500,
        "description": "One identifier: a recid (6004, recid:6004 or a portal record URL), a DOI (10.7483/OPENDATA.CMS.YLIC.86ZZ, doi:… or a doi.org URL), a CMS dataset path (/DoubleMuParked/Run2012B-22Jan2013-v1/AOD) or a documentation slug (cms-guide-docker or its portal URL)."
      }
    }
  },
  "required": [
    "ids"
  ],
  "additionalProperties": false
}
view source ↗

cern_opendata_list_files

open-world

List one record's files: its file indexes (groups of up to ~1,300 files) with their XRootD URI-list URLs, and per file the XRootD URI, HTTPS download URL, size, adler32 checksum and availability. Without index, returns the record's indexes and its regular files; with index, pages through that index's files. Files marked on demand sit on tape and must be requested on the record's portal page before download.

read
invocation
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "cern_opendata_list_files",
    "arguments": {
      "recid": "<recid>"
    }
  }
}
schema
{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "type": "object",
  "properties": {
    "recid": {
      "description": "Record id: up to 12 digits (6004), recid:6004, or a portal record URL. cern_opendata_search_records and cern_opendata_get_records return it.",
      "type": "string",
      "pattern": "^\\d{1,12}$"
    },
    "index": {
      "description": "A file index key from the indexes list, matched exactly (a .txt ending is read as .json). Omit to list the record's indexes and regular files.",
      "type": "string",
      "maxLength": 300
    },
    "cursor": {
      "description": "next_cursor from the previous page, unchanged, with the same recid and index. Omit for the first page.",
      "type": "string",
      "maxLength": 500
    },
    "limit": {
      "description": "Files per page, 1-500.",
      "default": 50,
      "type": "integer",
      "minimum": 1,
      "maximum": 500
    }
  },
  "required": [
    "recid",
    "limit"
  ],
  "additionalProperties": false
}
view source ↗

cern_opendata_get_analysis_env

open-world

Assemble what is needed to analyse a record: its container images, CMSSW release and global tag; the condition-data, VM and validated-run records for its run periods; example software that declares it works with the record; and quoted sections of the guides the record links (the first two are fetched). Container images, software and guide code are licensed separately from the CC0 data. Use cern_opendata_list_files for the record's files.

read
invocation
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "cern_opendata_get_analysis_env",
    "arguments": {
      "recid": "<recid>"
    }
  }
}
schema
{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "type": "object",
  "properties": {
    "recid": {
      "description": "Record id: up to 12 digits (6004), recid:6004, or a portal record URL. cern_opendata_search_records and cern_opendata_get_records return it.",
      "type": "string",
      "pattern": "^\\d{1,12}$"
    }
  },
  "required": [
    "recid"
  ],
  "additionalProperties": false
}
view source ↗

cern_opendata_get_validated_runs

open-world

Get a CMS validated-run (good-run) list, which certifies the luminosity sections that are good for physics in each run. Select it by a CMS collision dataset recid, a validated-run list recid, or a run period such as Run2012B (give exactly one of recid and run_period). Choose the full validation or the muons-only variant, and narrow to a run range; a dataset recid defaults the range to the first and last run the dataset lists. Returns the runs with their luminosity-section ranges and the list file download URL. CMS only.

read
invocation
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "cern_opendata_get_validated_runs",
    "arguments": {}
  }
}
schema
{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "type": "object",
  "properties": {
    "recid": {
      "description": "A CMS collision dataset recid (its linked list is used) or a validated-run list recid (used as named): up to 12 digits, recid:N or a portal record URL. Give this or run_period.",
      "type": "string",
      "pattern": "^\\d{1,12}$"
    },
    "run_period": {
      "description": "A CMS run period such as Run2012B, matched case-insensitively; 2012B also matches Run2012B. Give this or recid. cern_opendata_list_reference with topic run_periods lists them.",
      "type": "string",
      "maxLength": 40
    },
    "variant": {
      "description": "full (every detector certified) or muons_only (certified for muon physics). Omitted: a list recid is used as named; a dataset recid or run_period selects full.",
      "type": "string",
      "enum": [
        "full",
        "muons_only"
      ]
    },
    "run_min": {
      "description": "Lowest run to return, inclusive. With a dataset recid and neither bound set, run_min and run_max default to the first and last run the dataset lists; run_bounds echoes the range applied.",
      "type": "integer",
      "minimum": 1,
      "maximum": 9007199254740991
    },
    "run_max": {
      "description": "Highest run to return, inclusive. Defaults with run_min for a dataset recid.",
      "type": "integer",
      "minimum": 1,
      "maximum": 9007199254740991
    },
    "limit": {
      "description": "Runs to return, 1-2000.",
      "default": 200,
      "type": "integer",
      "minimum": 1,
      "maximum": 2000
    }
  },
  "required": [
    "limit"
  ],
  "additionalProperties": false
}
view source ↗

cern_opendata_search_trigger_paths

open-world

Look up CMS High-Level Trigger paths by exact name (HLT_IsoMu24) or prefix pattern (HLT_IsoMu*), optionally for one data-taking year. Each match is a per-year path record parsed into the first and last run the path was seen online, per-version run ranges, the L1 seed, and links to the HLT menu records. Covers CMS open data from 2010-2016. Prescale tables are not published.

read
invocation
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "cern_opendata_search_trigger_paths",
    "arguments": {
      "path": "<path>"
    }
  }
}
schema
{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "type": "object",
  "properties": {
    "path": {
      "description": "Trigger path name (HLT_IsoMu24) or prefix with one trailing wildcard (HLT_IsoMu*). Trimmed; HLT_ is added when missing and its case fixed. A trailing version suffix (_v3 or _v*) is dropped, since records list versions as V<n>.",
      "type": "string",
      "maxLength": 200,
      "pattern": "^HLT_[A-Za-z0-9_]+\\*?$"
    },
    "year": {
      "description": "Data-taking year, such as 2012. Trigger records cover 2010-2016.",
      "type": "integer",
      "minimum": 2000,
      "maximum": 2100
    },
    "limit": {
      "description": "Records per page, 1-50.",
      "default": 10,
      "type": "integer",
      "minimum": 1,
      "maximum": 50
    },
    "page": {
      "description": "Page number, from 1. page × limit may not exceed 10,000.",
      "default": 1,
      "type": "integer",
      "minimum": 1,
      "maximum": 9007199254740991
    }
  },
  "required": [
    "path",
    "limit",
    "page"
  ],
  "additionalProperties": false
}
view source ↗

cern_opendata_list_reference

Decode the vocabulary the other cern_opendata tools accept: experiments, record types, collision energies and types, file formats and data tiers, availability states, identifier forms, query syntax, licensing, and the CMS run periods that have validated-run lists. Static and offline; omit topic for every table.

read
invocation
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "cern_opendata_list_reference",
    "arguments": {}
  }
}
schema
{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "type": "object",
  "properties": {
    "topic": {
      "description": "One table to return: experiments, record_types, collision_energies, collision_types, file_types, availability, identifiers, query_syntax, licensing or run_periods. Omit for every table.",
      "type": "string",
      "enum": [
        "experiments",
        "record_types",
        "collision_energies",
        "collision_types",
        "file_types",
        "availability",
        "identifiers",
        "query_syntax",
        "licensing",
        "run_periods"
      ]
    }
  },
  "additionalProperties": false
}
view source ↗

Resources

1

One CERN Open Data Portal record's metadata by recid: description, run periods, collision and distribution details, related records, software-environment summary, license and citation. File lists are not included. Tool coverage: cern_opendata_get_records.

uri cern-opendata://record/{recid} mime application/json