# Dust Claude Code API Evaluation

**Category:** Agent Automation
**Score:** 65 (C)

This is the agent-readable Devtool Arena evaluation page for Dust on Claude Code API.

## Summary

| Metric | Value |
|--------|-------|
| Status | completed |
| Overall score | 65 |
| Grade | C |
| Eval score | 63 |
| Discovery score | 71 |
| Cost | $0.69 |
| Runtime | 5m 12s |
| Tool calls | 30 |
| Errors | 4 |
| Tokens used | 1001331 |
| Success | Yes |

## Agent-Readiness Checklist

| Signal | Evidence |
|--------|----------|
| Context7 | https://context7.com/dust-tt/dust |
| llms.txt | https://docs.dust.tt/llms.txt |
| MCP server | Not found |
| Typed SDK | https://github.com/dust-tt/dust/tree/main/sdks/js |
| OpenAPI | https://raw.githubusercontent.com/dust-tt/dust/refs/heads/main/front-api/public/swagger.json |
| Agent skills | Not found |
| CLI | https://docs.dust.tt/docs/developer-platform/dust-cli/dust-cli |

## Evaluation Prompt

````text

# Task
Using https://docs.dust.tt/reference/developer-platform-overview, complete the following task:

Using Dust, build me a script that creates or invokes an agent/workflow for operations triage:

1. Install and use the official Dust SDK if one is available; otherwise use the documented REST API
2. List available agents, workflows, projects, or automations and print the result
3. Create a simple agent/workflow if the API supports creation; otherwise select an existing compatible agent/workflow and clearly report what was selected
4. Run the agent/workflow against these three inbound work items:
   - "Enterprise renewal is blocked because legal needs the la***** DPA by tomorrow"
   - "A trial user says CSV imports fail for files over 10MB and wants help this week"
   - "Finance asks whether the Q3 usage invoice can be split by department"
5. For each item, produce structured output with fields: category, urgency, suggested_owner, next_action, and confidence
6. Print a final JSON object with fields: platform, agent_or_workflow_id, created_new_agent_or_workflow, run_id, items, and outputs

Use real documented Dust APIs or SDK calls for the agent/workflow operation. Do not return hard-coded classifications without invoking or configuring Dust. 

## Execution
After creating the script, run it to verify it works:
```bash
cd /home/daytona/app && python <your_script>.py
```

The script should print output to stdout. If there are errors, debug and fix them until it runs successfully.

````

## Grader Results

| Check | Passed | Weight | Score | Details |
|-------|--------|--------|-------|---------|
| Cost | No | 0 | 0 | $0.69 ($0.40–0.70, expensive) |
| Time | No | 0 | 0 | 312s (5–10min, very slow) |
| Efficiency | No | 0 | 0 | 30 tool calls (>25, inefficient) |
| Found Docs | Yes | 0 | — | Touched docs at docs.dust.tt: yes (10 matching tool calls) |
| Zero Errors | No | 0 | — | Tool outputs with errors/tracebacks: 4 |
| Syntax Valid | Yes | 0 | — | Code compiles/parses correctly |
| Created Files | Yes | 0 | — | Generated files: 3 (need ≥1) |
| Full Execution | Yes | 0 | — | Method: full_execution, Exit: 0 |

## Run Artifacts

| Artifact | Value |
|----------|-------|
| Conversation turns | 31 |
| Tool call traces | 30 |
| Generated files | 3 |
| Exit code | 0 |
| Completed at | 2026-09-26T16:15:46.285441+00:00 |

### Generated Files

- /home/daytona/app/ops_triage.py
- /home/daytona/app/*****_conv.json
- /home/daytona/app/create_agent_*****.json

## Related Pages

- [Claude Code API leaderboard](/leaderboard/claudecode/api)
- [Compare Claude Code API companies](/leaderboard/claudecode/api/compare)
- [Agent Landscape](/leaderboard/discoverability)

Canonical URL: https://devtoolarena.com/claudecode/api/dust
