Team Ai
Apppublic

EricMingle69/gaia-claude-code-agent

sourceHugging Facemitupdated 15d agoView on Hugging Face
0likes
App README

GAIA agent built on Claude Code (headless)

Final-assignment agent for the Hugging Face Agents Course (Unit 4, GAIA level-1 subset).

Design

  • —One question = one isolated headless Claude Code session (claude -p), so the agent loop, tool calling and context management come from Claude Code itself.
  • —Built-in tools: Bash, WebSearch, WebFetch, Read (images/PDFs), file editing.
  • —Local toolbox exposed through Bash: mlx_whisper speech-to-text, stockfish + python-chess, yt-dlp + ffmpeg frame extraction, youtube_transcript_api, pandas/openpyxl.
  • —System prompt (agent/system_prompt.md): the official GAIA answer-format rules, a verification method (primary sources, archived versions for dated facts, compute instead of estimate), and an integrity rule.
  • —No question-specific code, no hard-coded answers. The same prompt and toolbox are used for every question and every model.

Integrity

  • —The agent is instructed never to open GAIA answer keys, course leaderboards or other students' submissions.
  • —agent/audit.py scans every session transcript for actions that touched such sources (search queries, fetched URLs, shell commands) and lists them for manual review.
  • —Question texts and attachments are not redistributed here (GAIA terms).

Run

bash
python agent/gaia_agent.py --model <model-id> --effort high --workers 5
python agent/audit.py <model-id>__high
python agent/submit.py <model-id>__high

Requires a logged-in Claude Code CLI.

Result

claude-opus-5-5, effort high: 20/20, single submission, no prompt tuning or reruns. Median 2 turns and ~16 s per question.

Caveat: the GAIA validation set has been public since 2023. The audit shows the agent did not consult answer sources, but it cannot rule out that the model saw these questions during training, so this score shows the harness works end to end rather than measuring performance on unseen tasks.