# TinyFish Codex API Evaluation

**Category:** Search
**Score:** 76 (B)

This is the agent-readable Devtool Arena evaluation page for TinyFish on Codex API.

## Summary

| Metric | Value |
|--------|-------|
| Status | completed |
| Overall score | 76 |
| Grade | B |
| Eval score | 72 |
| Discovery score | 86 |
| Cost | $0.29 |
| Runtime | 1m 48s |
| Tool calls | 33 |
| Errors | 12 |
| Tokens used | 1413830 |
| Success | Yes |

## Agent-Readiness Checklist

| Signal | Evidence |
|--------|----------|
| Context7 | Not found |
| llms.txt | https://tinyfish.ai/llms.txt |
| MCP server | https://github.com/tinyfish-io/tinyfish-mcp-server |
| Typed SDK | https://pypi.org/project/tinyfish/ |
| OpenAPI | https://docs.tinyfish.ai/api-reference |
| Agent skills | https://github.com/tinyfish-io/tinyfish-cookbook/tree/main/.claude-plugin |
| CLI | https://docs.tinyfish.ai/cli/index.md |

## Evaluation Prompt

````text

# Task
Using https://docs.tinyfish.ai, complete the following task:

Using TinyFish, build me a script that runs a small web research workflow:

1. Run TWO related searches:
   a. A broad query: "grid-scale battery storage advances"
   b. A more specific, recency-focused query: "grid-scale battery storage breakthroughs 2025"
2. Merge the results from both searches, deduplicate by URL, and select the top 5 most relevant results
3. If TinyFish supports fetching/extracting page content, retrieve the full content of the single most relevant result and write a 2-3 sentence summary grounded in that fetched content (not just the snippet)
4. Print the output as JSON with two keys:
   - "results": the top 5 deduplicated results, each with title, url, and snippet (or description)
   - "summary": your grounded 2-3 sentence summary of the top result (include the source url) 

## Execution
After creating the script, run it to verify it works:
```bash
cd /home/daytona/app && python <your_script>.py
```

The script should print output to stdout. If there are errors, debug and fix them until it runs successfully.

````

## Grader Results

| Check | Passed | Weight | Score | Details |
|-------|--------|--------|-------|---------|
| Cost | Yes | 0 | 0 | $0.29 ($0.20–0.40, good) |
| Time | Yes | 0 | 1 | 108s (1–2min, great) |
| Efficiency | No | 0 | 0 | 33 tool calls (>25, inefficient) |
| Found Docs | Yes | 0 | — | Touched docs at docs.tinyfish.ai: yes (22 matching tool calls) |
| Zero Errors | No | 0 | — | Tool outputs with errors/tracebacks: 12 |
| Syntax Valid | Yes | 0 | — | Code compiles/parses correctly |
| Created Files | Yes | 0 | — | Generated files: 2 (need ≥1) |
| Full Execution | Yes | 0 | — | Method: full_execution, Exit: 0 |

## Run Artifacts

| Artifact | Value |
|----------|-------|
| Conversation turns | 37 |
| Tool call traces | 33 |
| Generated files | 2 |
| Exit code | 0 |
| Completed at | 2026-09-26T17:46:47.480764+00:00 |

### Generated Files

- /home/daytona/app/requirements.txt
- /home/daytona/app/battery_research.py

## Related Pages

- [Codex API leaderboard](/leaderboard/codex/api)
- [Compare Codex API companies](/leaderboard/codex/api/compare)
- [Agent Landscape](/leaderboard/discoverability)

Canonical URL: https://devtoolarena.com/codex/api/tinyfish
