# Tavily Codex MCP Evaluation

**Category:** Search
**Score:** 45 (D)

This is the agent-readable Devtool Arena evaluation page for Tavily on Codex MCP.

## Summary

| Metric | Value |
|--------|-------|
| Status | completed |
| Overall score | 45 |
| Grade | D |
| Eval score | 42 |
| Discovery score | 55 |
| Cost | $0.14 |
| Runtime | 2m 0s |
| Tool calls | 8 |
| Errors | 2 |
| Tokens used | 385393 |
| Success | No |

## Agent-Readiness Checklist

| Signal | Evidence |
|--------|----------|
| Context7 | Not found |
| llms.txt | Not found |
| MCP server | Not found |
| Typed SDK | Not found |
| OpenAPI | Not found |
| Agent skills | Not found |
| CLI | Not found |

## Evaluation Prompt

````text

# Task
Using the Tavily MCP, run a small research workflow:

1. Discover available MCP tools and their capabilities
2. Run TWO related searches to build a picture of a topic:
   a. A broad query: "grid-scale battery storage advances"
   b. A more specific, recency-focused query: "grid-scale battery storage breakthroughs 2025"
3. Merge the results from both searches, deduplicate by URL, and select the top 5 most relevant results
4. If the MCP exposes a content-fetch / page-extraction tool, use it to retrieve the full content of the single most relevant result, then write a 2-3 sentence summary that is grounded in that fetched content (not just the snippet)
5. Format the output as JSON with two keys:
   - "results": the top 5 deduplicated results, each with title, URL, and snippet/description
   - "summary": your grounded 2-3 sentence summary of the top result (include the source URL)

Use ONLY the MCP tools for all search and fetch operations.

You MUST use the Tavily MCP server to accomplish this task.
Use the MCP tools for the actual product operations. Do NOT use curl, fetch, raw HTTP requests, CLI commands, or SDK calls as substitutes.

Use the MCP tools' help/descriptions to understand what each tool does.
If you need additional context, check the documentation: https://docs.tavily.com/documentation/mcp

## Setup
Do NOT assume the Tavily MCP server is already installed or already running in this sandbox.
Review the installation and authentication details below before starting the task.

### Installation
Transport: `stdio`
Codex is configured to launch the `tavily` MCP server over stdio using: `npx -y tavily-mcp@la*****`
Stored MCP package hint: `tavily-mcp`
No separate setup command is required before startup unless the docs indicate one.
Bootstrap notes: Seeded from locally validated MCP config in backend/mcp-claude-working.json on 2026-04-13. Uses the la***** Tavily MCP package via npx.

### Authentication
The following environment variables are already set in this environment: `TAVILY_API_KEY`
The MCP server should read them automatically when it starts.

## Execution
1. The harness has already written the launch configuration for `tavily`. Start with the task directly; do not rewrite agent settings, `.mcp-runtime.json`, or build a separate MCP client unless the connected server still fails after a normal retry.
2. You do not need to enumerate every MCP tool up front; inspect only the relevant tools you need. Agent-facing MCP tool names may be namespaced like `mcp__tavily__<tool_name>`.
3. Complete the task described above using MCP tools for the actual product operations.
4. If the server fails to start or authenticate, debug the documented setup and retry. Do not fall back to raw HTTP, SDK, CLI, or manual JSON-RPC/MCP calls as a substitute for exposed MCP tools.

Work in the /home/daytona/app directory.
If you use Bash during setup or debugging, limit it to installation/configuration work and then return to MCP tools for the task itself.
If there are errors, debug and fix them until the task runs successfully.

## Summary
When you are done, write a JSON summary file to /home/daytona/app/results.json with this structure:
{
  "operations": [
    {"step": 1, "description": "What you did", "mcp_tool": "tool_name", "success": true/false, "output_summary": "brief result"},
    ...
  ],
  "overall_success": true/false,
  "notes": ["any relevant notes about the execution"]
}

````

## Grader Results

| Check | Passed | Weight | Score | Details |
|-------|--------|--------|-------|---------|
| Task | No | — | — | — |
| Llms Txt | Yes | — | 10 | — |
| Tool Count | Yes | — | 10 | — |
| Auth Method | Yes | — | 15 | — |
| Official Mcp | Yes | — | 5 | — |
| Output Schema | No | — | 0 | — |
| Tool Stability | Yes | — | 5 | — |
| Tool Description Length | Yes | — | 10 | — |

## Run Artifacts

| Artifact | Value |
|----------|-------|
| Conversation turns | 14 |
| Tool call traces | 8 |
| Generated files | 1 |
| Exit code | — |
| Completed at | 2026-09-27T15:59:25.645821+00:00 |

### Generated Files

- /home/daytona/app/results.json

## Related Pages

- [Codex MCP leaderboard](/leaderboard/codex/mcp)
- [Compare Codex MCP companies](/leaderboard/codex/mcp/compare)
- [Agent Landscape](/leaderboard/discoverability)

Canonical URL: https://devtoolarena.com/codex/mcp/tavily
