# Tenki Codex API Evaluation

**Category:** Sandboxes
**Score:** 75 (B)

This is the agent-readable Devtool Arena evaluation page for Tenki on Codex API.

## Summary

| Metric | Value |
|--------|-------|
| Status | completed |
| Overall score | 75 |
| Grade | B |
| Eval score | 77 |
| Discovery score | 71 |
| Cost | $0.15 |
| Runtime | 1m 15s |
| Tool calls | 18 |
| Errors | 4 |
| Tokens used | 551050 |
| Success | Yes |

## Agent-Readiness Checklist

| Signal | Evidence |
|--------|----------|
| Context7 | Not found |
| llms.txt | https://tenki.cloud/llms.txt |
| MCP server | https://tenki.cloud/docs/sandbox/mcp |
| Typed SDK | https://registry.npmjs.org/@tenkicloud/sandbox |
| OpenAPI | Not found |
| Agent skills | https://tenki.cloud/.well-known/agent-skills/index.json |
| CLI | https://tenki.cloud/docs/sandbox/cli |

## Evaluation Prompt

````text

# Task
Using https://tenki.cloud/docs/sandbox/sdk, complete the following task:

Using Tenki, build me a script that runs a multi-step computation inside a Tenki sandbox:
        
1. Use the Tenki Python SDK to create/connect to a sandbox
2. Inside the sandbox, execute the following pipeline:
   a. Install the `sympy` package using pip
   b. Run a Python script that:
      - Computes the 10th Fibonacci number iteratively
      - Uses sympy to check if that number is prime
      - Prints the result as JSON: {"fibonacci_10": <number>, "is_prime": <boolean>}
3. Capture the sandbox execution output and print it to stdout

All computation must run inside the Tenki sandbox, not in your local Python process. Do not use subprocess — use the Tenki SDK's execution primitives. 

## Execution
After creating the script, run it to verify it works:
```bash
cd /home/daytona/app && python <your_script>.py
```

The script should print output to stdout. If there are errors, debug and fix them until it runs successfully.

````

## Grader Results

| Check | Passed | Weight | Score | Details |
|-------|--------|--------|-------|---------|
| Cost | Yes | 0 | 1 | $0.15 ($0.10–0.20, great) |
| Time | Yes | 0 | 1 | 75s (1–2min, great) |
| Efficiency | Yes | 0 | 0 | 18 tool calls (≤20, acceptable) |
| Found Docs | Yes | 0 | — | Touched docs at tenki.cloud: yes (10 matching tool calls) |
| Zero Errors | No | 0 | — | Tool outputs with errors/tracebacks: 4 |
| Syntax Valid | Yes | 0 | — | Code compiles/parses correctly |
| Created Files | Yes | 0 | — | Generated files: 2 (need ≥1) |
| Full Execution | Yes | 0 | — | Method: full_execution, Exit: 0 |

## Run Artifacts

| Artifact | Value |
|----------|-------|
| Conversation turns | 22 |
| Tool call traces | 18 |
| Generated files | 2 |
| Exit code | 0 |
| Completed at | 2026-09-26T17:46:23.853131+00:00 |

### Generated Files

- /home/daytona/app/requirements.txt
- /home/daytona/app/tenki_fibonacci.py

## Related Pages

- [Codex API leaderboard](/leaderboard/codex/api)
- [Compare Codex API companies](/leaderboard/codex/api/compare)
- [Agent Landscape](/leaderboard/discoverability)

Canonical URL: https://devtoolarena.com/codex/api/tenki
