textFujitsu-FRE/MAPSagentsbenchmarkmultilingualevaluationsecuritygaiaswe-benchmathagentic-aisafety

MAPS Multilingual Agentic AI Benchmark

Free

Open dataset

Sample structure: 82.1 / 100
1 download links issued
Seller: DataBazaar
Sign up to download

Already have an account? Log in

Agent? Connect your account →

Category
Text
Records
9,680 rows
Format
PARQUET
Update Frequency
Not documented
Collection Method
auto_imported_huggingface_federated
PII
No flagged field names; not a privacy audit
File Size
~8.92 MB
Download links issued
1

Source, license and coverage

Supplier documentation. These claims are separate from the automated sample score. A listing edit date is not a data freshness date.

License
cc-by-4.0
Source / creator
Fujitsu-FRE/MAPS
Collection method
The authors curated 405 performance tasks from established agentic/reasoning benchmarks (GAIA, SWE-bench, MATH) and 400 tasks from the Agent Security Benchmark, then produced multilingual versions across 11 languages to enable cross-lingual evaluation of both capability and security. Translation/localization methodology is described in the accompanying paper; tasks preserve original answer keys for evaluation parity with the English source benchmarks.
Coverage start
Not documented
Coverage end
Not documented
Data last updated
Not documented
Update schedule
Not documented

Source documentation ↗

License terms ↗

Multilingual benchmarks built via translation can carry translation artifacts and may not capture culturally-specific reasoning. Coverage is limited to 11 languages, weighted toward high-resource languages. Because tasks derive from public benchmarks (GAIA, SWE-bench, MATH), there is meaningful overlap with widely-used training corpora — leakage risk for any model trained on web-scale data. The 805-task scale is modest and may have high evaluation variance. Source does not exhaustively document per-language quality assurance; buyers should validate empirically for their use case.

Sample structure score: 82.1 / 100

This automated check describes the inspected sample, not factual accuracy, legal rights, representativeness, or the quality of the entire dataset. It is not a customer rating.

Assessed 10 sample records (JSON) on 2026-10-09. All records in the provided sample were checked.

CheckPointsEvidence
Populated cells35 / 5042 of 60 top-level cells contain a value. Null, absent and blank values count as missing; zero and false count as populated.
Consistent value types27.1 / 3038 of 42 populated cells match their column's most common observed type. Types are inferred, not checked against real-world truth.
Consistent record shape20 / 2010 of 10 records have the expected fields. CSV/TSV use the header width; JSON uses the union of observed keys.
Field-level findings and improvements

Check missing cells and mixed types below. Document intentional missing values or mixed types in your field descriptions. Do not fill legitimate unknowns with invented values just to increase this score.

FieldMissing cellsMost common typeOther populated types
task_id0 / 10string0 / 10
Question0 / 10string0 / 10
Level0 / 10number0 / 10
Final answer0 / 10number4 / 10
file_name9 / 10string0 / 1
file_path9 / 10string0 / 1
How the score is calculated, its limitations, and how to correct an assessment →

About this data

Benchmark combining performance and security evaluation tasks across 11 languages, integrating GAIA, SWE-bench, and MATH datasets with agentic security assessments.

Retrieve with your agent or Python

Create an account and configure DATABAZAAR_API_KEY. This example retrieves free or already purchased data; it never makes a purchase. For a multi-file dataset, choose a file index from the manifest.

Download the Python example
python3 retrieve-dataset.py 5deeba74-7914-4cce-81f5-ad60f86f7a74 --output dataset.bin
Full supplier documentation
## Overview MAPS (Multilingual Agentic AI Benchmark for Global Agent Performance and Security) is a benchmark of 805 tasks designed to evaluate agentic AI systems across 11 languages and diverse tasks. It combines 405 performance-oriented tasks drawn from GAIA, SWE-bench, and MATH with 400 tasks from the Agent Security Benchmark, enabling systematic analysis of agent behavior under multilingual conditions for both capability and safety dimensions. Text modality, 1K–10K size range. ## Schema Schema follows standard agent-eval task records. Typical fields include: - `task_id` — string — unique identifier for each task - `language` — string — ISO language code (ar, en, ja, es, ko, hi, ru, he, pt, de, it) - `source_dataset` — string — origin benchmark (GAIA, SWE-bench, MATH, Agent Security Benchmark) - `category` — string — performance vs. security classification - `question` / `prompt` — string — the task instruction in the target language - `answer` / `expected_output` — string — ground-truth answer or expected behavior - `metadata` — object — task-specific fields (difficulty, tools required, threat model, etc.) Exact columns vary per source subset; consult the HF dataset viewer for per-split fields. ## Sources - Fujitsu-FRE/MAPS on Hugging Face — https://huggingface.co/datasets/Fujitsu-FRE/MAPS — license: CC-BY-4.0 - Underlying source benchmarks: GAIA (arXiv:2311.12983), SWE-bench (arXiv:2310.06770), MATH (arXiv:2103.03874), Agent Security Benchmark (arXiv:2410.02644) ## Methodology The authors curated 405 performance tasks from established agentic/reasoning benchmarks (GAIA, SWE-bench, MATH) and 400 tasks from the Agent Security Benchmark, then produced multilingual versions across 11 languages to enable cross-lingual evaluation of both capability and security. Translation/localization methodology is described in the accompanying paper; tasks preserve original answer keys for evaluation parity with the English source benchmarks. ## Known gaps & limitations Multilingual benchmarks built via translation can carry translation artifacts and may not capture culturally-specific reasoning. Coverage is limited to 11 languages, weighted toward high-resource languages. Because tasks derive from public benchmarks (GAIA, SWE-bench, MATH), there is meaningful overlap with widely-used training corpora — leakage risk for any model trained on web-scale data. The 805-task scale is modest and may have high evaluation variance. Source does not exhaustively document per-language quality assurance; buyers should validate empirically for their use case. ## Intended use & out-of-scope - IS for: multilingual evaluation of agentic AI systems, cross-lingual performance/security comparisons, regression testing of agent frameworks across languages. - NOT for: training data (high contamination risk against public benchmarks), production safety certification, or as a substitute for native-language adversarial red-teaming. _Federated dataset: 49 parquet shards, 8.9 MB total. Queries and downloads stream through the DataBazaar API._ Original supplier listing: MAPS: Multilingual Agentic AI Benchmark (Performance & Security) 805-task multilingual benchmark for evaluating agentic AI across 11 languages, combining performance tasks (GAIA, SWE-bench, MATH) with Agent Security Benchmark tasks. CC-BY-4.0.

Schema

NameTypeDescription
task_idVARCHARUUID string uniquely identifying each task in the benchmark
QuestionVARCHARTask instruction or query in target language requiring agent completion
LevelVARCHARDifficulty rating (1-3) indicating task complexity
Final answerVARCHARGround-truth answer or expected output for task evaluation
file_nameVARCHAROriginal filename of the task source document or resource
file_pathVARCHARFile path or URL reference to the source document or resource

Sample Data

Preview a sample of the data before downloading.

Public sample only. Sign in to retrieve the full dataset, including free datasets.

For AI Agents

Via MCP Server
# 1. Add to your agent's MCP config (claude_desktop_config.json or similar):
{
  "mcpServers": {
    "databazaar": { "command": "npx", "args": ["databazaar-mcp"] }
  }
}

# 2. Your agent can then call:
search_datasets({ query: "MAPS Multilingual Agentic AI B" })
// Found: 5deeba74-7914-4cce-81f5-ad60f86f7a74
get_download_url({ dataset_id: "5deeba74-7914-4cce-81f5-ad60f86f7a74" })  // free — sign in with MCP OAuth first
Via REST API
# Free dataset — sign in or use your account API key:
curl https://api.databazaar.io/datasets/5deeba74-7914-4cce-81f5-ad60f86f7a74/download-url -H "Authorization: Bearer $DATABAZAAR_API_KEY"