# Letta Codex API Evaluation

**Category:** Agent Memory
**Score:** 73 (C)

This is the agent-readable Devtool Arena evaluation page for Letta on Codex API.

## Summary

| Metric | Value |
|--------|-------|
| Status | completed |
| Overall score | 73 |
| Grade | C |
| Eval score | 68 |
| Discovery score | 86 |
| Cost | $0.34 |
| Runtime | 2m 14s |
| Tool calls | 34 |
| Errors | 2 |
| Tokens used | 974003 |
| Success | Yes |

## Agent-Readiness Checklist

| Signal | Evidence |
|--------|----------|
| Context7 | https://context7.com/letta-ai/letta |
| llms.txt | https://docs.letta.com/llms.txt |
| MCP server | https://docs.letta.com/platform/hosted-mcp |
| Typed SDK | https://docs.letta.com/v1-sdk/client-sdks |
| OpenAPI | Not found |
| Agent skills | https://github.com/letta-ai/skills |
| CLI | https://docs.letta.com/quickstart |

## Evaluation Prompt

````text

# Task
Using https://docs.letta.com/v1-sdk/client-sdks, complete the following task:

Using the Letta Python SDK, build me a long-term memory layer for an AI assistant and prove that, after the user corrects a fact in a later conversation, the assistant answers with the user's CURRENT state:

1. Install the official Letta Python SDK (pip install) and initialize the client/memory store for a single user id "leaderboard-user".

2. Simulate CONVERSATION 1 by storing these facts for the user:
   - "My name is Sam and I live in Berlin."
   - "I am allergic to peanuts."
   - "I prefer window seats when I fly."

3. Simulate CONVERSATION 2 (a later session for the SAME user): the user says "I moved from Berlin to Amsterdam last month." Record this correction using whatever mechanism Letta provides. Any idiomatic Letta approach is acceptable — for example automatic reconciliation on add, an explicit update call, deleting/replacing the old fact, editing a memory block, or re-processing the memory graph. Do not force a specific mechanism; use the one Letta is designed for.

4. Retrieve the memories relevant to this query: "Where does the user live and is there anything I should know before booking a restaurant for them?"

5. Print the result as JSON with these keys:
   - "query": the query string
   - "retrieved_memories": the list of memory strings Letta returned for the query
   - "current_city": the user's current city, as reflected by what Letta returned
   - "dietary_note": any dietary constraint, as reflected by what Letta returned

The correction must be recorded through Letta itself so that Letta's own retrieval reflects the CURRENT state (Amsterdam, not Berlin) — for example by returning only the current location, or by marking the Berlin fact as outdated/superseded. It is NOT acceptable to store two contradictory facts and then have your own code (or a separate LLM call) pick the newer one; the resolution must come from Letta. Use Letta's own memory storage/retrieval primitives for steps 2-4 — do NOT keep the memories only in a local Python list, and do NOT make raw HTTP/REST calls. 

## Execution
After creating the script, run it to verify it works:
```bash
cd /home/daytona/app && python <your_script>.py
```

The script should print output to stdout. If there are errors, debug and fix them until it runs successfully.

````

## Grader Results

| Check | Passed | Weight | Score | Details |
|-------|--------|--------|-------|---------|
| Cost | Yes | 0 | 0 | $0.34 ($0.20–0.40, good) |
| Time | Yes | 0 | 0 | 134s (2–3min, good) |
| Efficiency | No | 0 | 0 | 34 tool calls (>25, inefficient) |
| Found Docs | Yes | 0 | — | Touched docs at docs.letta.com: yes (25 matching tool calls) |
| Zero Errors | No | 0 | — | Tool outputs with errors/tracebacks: 2 |
| Syntax Valid | Yes | 0 | — | Code compiles/parses correctly |
| Created Files | Yes | 0 | — | Generated files: 1 (need ≥1) |
| Full Execution | Yes | 0 | — | Method: full_execution, Exit: 0 |

## Run Artifacts

| Artifact | Value |
|----------|-------|
| Conversation turns | 40 |
| Tool call traces | 34 |
| Generated files | 1 |
| Exit code | 0 |
| Completed at | 2026-09-27T13:32:01.395407+00:00 |

### Generated Files

- /home/daytona/app/long_term_memory.py

## Related Pages

- [Codex API leaderboard](/leaderboard/codex/api)
- [Compare Codex API companies](/leaderboard/codex/api/compare)
- [Agent Landscape](/leaderboard/discoverability)

Canonical URL: https://devtoolarena.com/codex/api/letta
