# Steel Claude Code CLI Evaluation

**Category:** Browser
**Score:** 80 (B)

This is the agent-readable Devtool Arena evaluation page for Steel on Claude Code CLI.

## Summary

| Metric | Value |
|--------|-------|
| Status | completed |
| Overall score | 80 |
| Grade | B |
| Eval score | 75 |
| Discovery score | 93 |
| Cost | $0.15 |
| Runtime | 1m 32s |
| Tool calls | 15 |
| Errors | 2 |
| Tokens used | 409711 |
| Success | Yes |

## Agent-Readiness Checklist

| Signal | Evidence |
|--------|----------|
| Context7 | Not found |
| llms.txt | Not found |
| MCP server | Not found |
| Typed SDK | Not found |
| OpenAPI | Not found |
| Agent skills | Not found |
| CLI | Not found |

## Evaluation Prompt

````text

# Task
Using the steel CLI (Steel), automate a browser session:

1. Open https://example.com in a browser session
2. Read the page title and main heading text using CLI commands
3. Capture a screenshot to /home/daytona/app/browserbase-example.png
4. Print a JSON summary with url, title, heading_text, and screenshot_path

Use the CLI directly, not Python SDK imports or raw HTTP calls.

You MUST use the `steel` CLI to accomplish this task.
Do NOT write Python scripts that import the SDK — use CLI commands via the terminal.

If you need help with CLI commands, check the documentation: https://docs.steel.dev/overview/steel-cli
You can also use `steel --help` and `steel <command> --help` to discover available commands.

## Setup
You need to install and authenticate the `steel` CLI yourself.

### Installation
Run: `npm install -g @steel-dev/cli`

### Authentication
The following environment variables are already set in this environment: `STEEL_API_KEY`
The CLI should pick these up automatically, or pass them to commands as needed.

## Execution
1. Install the `steel` CLI using the command above
2. Verify the installation: `steel --help`
3. Authenticate using the credentials above
4. Perform the task described above

Work in the /home/daytona/app directory. Run commands and print output to stdout.
If there are errors, debug and fix them until the task runs successfully.

## Summary
When you are done, write a JSON summary file to /home/daytona/app/results.json with this structure:
{
  "operations": [
    {"step": 1, "description": "What you did", "command": "the CLI command", "success": true/false, "output_summary": "brief result"},
    ...
  ],
  "overall_success": true/false,
  "notes": ["any relevant notes about the execution"]
}

````

## Grader Results

| Check | Passed | Weight | Score | Details |
|-------|--------|--------|-------|---------|
| Stage 2 Auth | Yes | — | 25 | — |
| Stage 3 Task | Yes | — | — | — |
| Stage 1 Install | Yes | — | 5 | — |
| Stage 5 Llms Txt | Yes | — | 5 | — |
| Stage 0 Cli Exists | Yes | — | 5 | — |
| Stage 4 Json Output | Yes | — | 15 | — |
| Stage 6 Agent Skill | Yes | — | 20 | — |
| Stage 3 Non Interactive | Yes | — | 18 | — |

## Run Artifacts

| Artifact | Value |
|----------|-------|
| Conversation turns | 21 |
| Tool call traces | 15 |
| Generated files | 1 |
| Exit code | — |
| Completed at | 2026-09-27T14:04:51.369168+00:00 |

### Generated Files

- /home/daytona/app/results.json

## Related Pages

- [Claude Code CLI leaderboard](/leaderboard/claudecode/cli)
- [Compare Claude Code CLI companies](/leaderboard/claudecode/cli/compare)
- [Agent Landscape](/leaderboard/discoverability)

Canonical URL: https://devtoolarena.com/claudecode/cli/steel
