# Datadog Claude Code CLI Evaluation

**Category:** Observability
**Score:** 63 (C)

This is the agent-readable Devtool Arena evaluation page for Datadog on Claude Code CLI.

## Summary

| Metric | Value |
|--------|-------|
| Status | completed |
| Overall score | 63 |
| Grade | C |
| Eval score | 68 |
| Discovery score | 53 |
| Cost | $0.25 |
| Runtime | 2m 36s |
| Tool calls | 21 |
| Errors | 7 |
| Tokens used | 693865 |
| Success | Yes |

## Agent-Readiness Checklist

| Signal | Evidence |
|--------|----------|
| Context7 | Not found |
| llms.txt | Not found |
| MCP server | Not found |
| Typed SDK | Not found |
| OpenAPI | Not found |
| Agent skills | Not found |
| CLI | Not found |

## Evaluation Prompt

````text

# Task
Using the pup CLI (Datadog), perform observability operations:

1. Submit a custom metric called "*****.leaderboard.cli_score" with value 42 and tags "env:*****,leaderboard:sapient"
2. Create a ***** event with title "CLI Leaderboard Test" and text "Integration ***** event from [REDACTED:ORG] CLI Leaderboard"
3. Print the API response or status for each operation

You MUST use the `pup` CLI to accomplish this task.
Do NOT write Python scripts that import the SDK — use CLI commands via the terminal.

If you need help with CLI commands, check the documentation: https://docs.datadoghq.com/cli/
You can also use `pup --help` and `pup <command> --help` to discover available commands.

## Setup
You need to install and authenticate the `pup` CLI yourself.

### Installation
Run: `brew tap datadog-labs/pack`
Run: `brew install pup`

### Authentication
The following environment variables are already set in this environment: `DD_API_KEY`, `DD_APP_KEY`, `DATADOG_API_KEY`, `DATADOG_APP_KEY`
The CLI should pick these up automatically, or pass them to commands as needed.

## Execution
1. Install the `pup` CLI using the command above
2. Verify the installation: `pup --help`
3. Authenticate using the credentials above
4. Perform the task described above

Work in the /home/daytona/app directory. Run commands and print output to stdout.
If there are errors, debug and fix them until the task runs successfully.

## Summary
When you are done, write a JSON summary file to /home/daytona/app/results.json with this structure:
{
  "operations": [
    {"step": 1, "description": "What you did", "command": "the CLI command", "success": true/false, "output_summary": "brief result"},
    ...
  ],
  "overall_success": true/false,
  "notes": ["any relevant notes about the execution"]
}

````

## Grader Results

| Check | Passed | Weight | Score | Details |
|-------|--------|--------|-------|---------|
| Stage 2 Auth | Yes | — | 3 | — |
| Stage 3 Task | Yes | — | — | — |
| Stage 1 Install | Yes | — | 5 | — |
| Stage 5 Llms Txt | Yes | — | 5 | — |
| Stage 0 Cli Exists | Yes | — | 5 | — |
| Stage 4 Json Output | Yes | — | 15 | — |
| Stage 6 Agent Skill | Yes | — | 20 | — |
| Stage 3 Non Interactive | No | — | 0 | — |

## Run Artifacts

| Artifact | Value |
|----------|-------|
| Conversation turns | 30 |
| Tool call traces | 21 |
| Generated files | 2 |
| Exit code | — |
| Completed at | 2026-09-27T14:04:24.985775+00:00 |

### Generated Files

- /home/daytona/app/results.json
- /home/daytona/app/metric_payload.json

## Related Pages

- [Claude Code CLI leaderboard](/leaderboard/claudecode/cli)
- [Compare Claude Code CLI companies](/leaderboard/claudecode/cli/compare)
- [Agent Landscape](/leaderboard/discoverability)

Canonical URL: https://devtoolarena.com/claudecode/cli/datadog
