Jess Wang13 minTransactional email delivers password resets, receipts, account alerts, and onboarding messages. At scale, teams rely on batches to send many personalized emails at once and templates to keep high-volume messages consistent. A pipeline failure can leave emails unsent or send an invoice or account detail to the wrong recipient.
AI agents and MCP servers can automate parts of that workflow. Transactional messages can contain password-reset links, invoice details, account information, and personal recipient data. This evaluation compares the Resend MCP and Postmark MCP across common tasks, direct acceptance timing, and record availability for follow-up lookup.
Full task evaluation: Codex completes nine email-task types through each MCP and returns a structured response. This measures workflow correctness and surfaces behavior in the traces.
Direct benchmark: The task evaluation includes agent reasoning and tool-use time, so it cannot isolate MCP latency. The direct benchmark uses a non-agent client. It tests one 10-email batch at a time, then low- and high-concurrency bursts of two and five batches. Each measurement ends when the tool returns accepted IDs.
Feature comparison: Some features could not be included in a fair head-to-head task because only one MCP exposed them. The feature comparison makes the scope of the two tools clear.
Source experiments and traces are available in the Braintrust project.
The following terms clarify the email pipeline used throughout the evaluation.
| Term | Meaning |
|---|---|
| Accepted ID | The provider's canonical identifier for one accepted email. Codex receives it after the provider accepts the send, uses that same ID for lookup and retries, and includes it in its final report so the harness can identify the exact email. Acceptance does not mean the email is in the recipient's inbox. |
| Provider record | The provider's stored record for that accepted ID. It includes the recipient, subject, content, tag, and current provider state. |
| Status visibility | How long after acceptance it takes for the same ID to return a usable record through the MCP lookup tool. Before then, the tool can return not found, unavailable, an error, or a response without a matching ID and usable state. |
| Batch | One native tool call containing separate messages for several recipients. Each recipient should receive a separate provider ID. A three-recipient batch should return three IDs. |
| Email unit | One required recipient-specific email. A 10-recipient batch is one task and one MCP call, but 10 email units. |
Codex receives an email task from the dataset and calls either the Resend MCP or Postmark MCP. The provider accepts the request and returns an email ID to Codex. Codex then calls the MCP again with that ID to retrieve the provider record. Finally, Codex returns structured JSON to the eval harness with the recipient, ID, and observed status.
The dataset contains nine task buckets with five variants per bucket.
not found response causes Codex to retry only the lookup and not create a duplicate send.| Task shape | One controlled example | What a failure would look like |
|---|---|---|
| Content / serialization direct send | Send Ada ExampleCo café update ☕ — 01. The body starts Hi Zoë 👋, and includes https://example.test/reports/cafe?locale=fr&source=v2. The HTML link is labelled Open report. Other variants use curly quotes, &, <v2.1>, Japanese text, apostrophes, and multiline copy. | café becomes café. Emoji or curly quotes disappear. <v2.1> is interpreted as markup rather than text, a blank line disappears, or the link loses &source=v2. |
| Shuffled five-recipient batch | The requested order is Sam, Lin, Ada, Mika, then Rin. Sam receives ExampleCo mapping MAP-5-01-03 for Sam, amount $34.25, and a Review MAP-5-01-03 link with recipient=3. Every other recipient has different copy and a different link. | All five emails are sent, but Sam receives Lin's amount or link. Two payloads are swapped, one recipient is missing, or the agent sends five individual messages instead of one native batch. |
| Two-recipient welcome batch | In one batch, Ada receives Welcome to ExampleCo, Ada and https://example.test/dashboard/ada?cohort=1. Lin receives the Lin subject and https://example.test/dashboard/lin?cohort=1. | Lin gets Ada's dashboard URL or name. Both recipients receive identical content, the agent combines the recipients in one email, or it uses two ordinary sends rather than one native batch. |
The evaluation uses seven deterministic scorers to test different parts of this pipeline.
| Scorer | Question |
|---|---|
provider_send_accepted | Did the provider accept every required email and return an ID? |
agent_status_retrieved | Did Codex itself get a usable record for every ID before it ended? A usable record has the matching provider ID and a real state, such as Sent or delivered. |
direct_payload_fidelity | Did Codex pass the requested subject, text, HTML/link, and tag to the MCP? |
provider_payload_fidelity | Did the provider record retain the required sender, content, and tag? |
batch_semantics | Did Codex use one native batch with distinct messages, rather than a sequence of sends? |
recipient_safety | Did every message go only to the intended recipient, with no CC/BCC or cross-recipient mix-up? |
final_report_identity | Did the final report include one correct recipient and canonical provider ID for every expected email? |
This diagram shows how each deterministic scorer verifies a specific behavior in the email workflow.
The task evaluation ran twice with the same dataset and configuration. The results below combine both runs.
| Scorer | Postmark | Resend |
|---|---|---|
provider_send_accepted | 240/240 | 240/240 |
agent_status_retrieved | 240/240 | 240/240 |
batch_semantics | 40/40 batch task rows | 40/40 batch task rows |
final_report_identity | 90/90 | 90/90 |
direct_payload_fidelity | 950/950 requested send-field comparisons | 950/950 requested send-field comparisons |
provider_payload_fidelity | 1,190/1,190 requested stored-field comparisons | 1,190/1,190 requested stored-field comparisons |
recipient_safety | 90/90 task rows | 90/90 task rows |
These deterministic checks show that both MCPs completed the tasks correctly. They do not identify a clear winner between Resend and Postmark.
First-lookup success differed between the two providers. It measures whether the first MCP lookup after the provider accepted an email returned a usable provider record rather than a temporary not found response.
| Measure | Postmark | Resend |
|---|---|---|
| First same-ID lookup usable | 166/240 = 69.2% | 239/240 = 99.6% |
The difference in first-lookup success was 30.4 percentage points. Postmark often accepted the email before the agent could retrieve its record to verify content and status. This matters for two reasons. The agent must retry the lookup, which adds latency and cost. An agent that interprets the missing record as a failed send could incorrectly report failure or send a duplicate email.
For Postmark, the first lookup for 74/240 email units returned an unavailable result and needed a retry, which led to 98 extra Postmark lookup calls.
If we group email units by the number in each send-tool call, larger batches had fewer Postmark first-lookup misses, probably because they give records more time to become available before Codex looks them up. In one 10-recipient Postmark trace, the batch was accepted at 52.12s and the first lookup began 15.04s later, giving the provider record time to appear.
| Email units in the original send call | Postmark email units needing a retry | Postmark median time to first successful lookup | Resend email units needing a retry |
|---|---|---|---|
| 1 | 42/50 (84%) | 9.68s | 0/50 |
| 2 | 20/40 (50%) | 12.17s | 0/40 |
| 5 | 12/50 (24%) | 10.81s | 0/50 |
| 10 | 0/100 (0%) | 15.04s | 1/100 |
The next question was whether the delay came from the MCP wrapper or the provider. A direct REST control called each provider's API with no Codex agent or MCP server. It used one task from each of the nine task buckets.
| Measure | Postmark | Resend |
|---|---|---|
| Accepted units | 24/24 | 24/24 |
| Median REST status-visibility time | 11.79s | 1.27s |
The delay before the first readable provider record remained even after the MCP and agent were removed. This indicates that the observed delay comes from the Postmark provider rather than the Postmark MCP.
To calculate the estimated cost, we used the traces' token counts. The estimate uses the current GPT-5.6 Sol prices of $4 per million input tokens, $0.40 per million cached input tokens, and $20 per million output tokens (GPT-5.6 Sol pricing). It excludes Postmark and Resend plan charges.
| Measure | Postmark total / avg | Resend total / avg |
|---|---|---|
| Prompt tokens | 10.64M | 9.55M |
| Cached prompt tokens | 9.43M | 8.45M |
| Output tokens | 92.8K | 100.1K |
| Reconstructed LLM cost | $10.45 | $9.80 |
| Reconstructed LLM cost per task | $0.116 | $0.109 |
Based on the rounded totals above, Postmark's estimated LLM cost was about $0.65 higher across the entire comparison, or about $0.007 more per task. This is probably because Postmark made 68 more model calls, consistent with extra lookup retries caused by the delay between email acceptance and provider-record availability.
The second part of this head-to-head comparison calls each MCP with no agent involved, which isolates MCP behavior.
The benchmark used three test conditions.
The single-batch test is the baseline latency condition. It measures one batch-acceptance round trip without competing calls. The low- and high-concurrency tests measure burst behavior, testing whether the MCP and provider stay responsive when two or five clients submit batches at the same time.
| Test condition | Result |
|---|---|
| Single batch | Postmark returned accepted IDs faster in all 30 matched 10-email batch calls, averaging 0.156 seconds sooner (95% CI, 0.095–0.218 seconds). |
| Low concurrency | Postmark was faster in all 25 matched waves, averaging 0.219 seconds sooner (95% CI, 0.072–0.366 seconds). |
| High concurrency | Postmark was faster in all 25 matched burst waves, averaging 0.173 seconds sooner (95% CI, 0.154–0.192 seconds). |
Postmark returned accepted IDs sooner after a direct batch request. Resend made those accepted email records available sooner when Codex immediately tried to look them up, which reduced retry work.
Some features could not be evaluated head to head because only one MCP exposed them. The differences were substantial enough to document separately.
| Feature | Postmark MCP | Resend MCP | Why it matters |
|---|---|---|---|
| Direct transactional send | Yes | Yes | Shared workflow tested head to head. |
| Native direct-content batch | Yes, up to 500 provider items | Yes, up to 100 items | Shared workflow tested up to 50 items. |
| Return all IDs from a 50-item successful batch | No. Only the first 10 IDs were shown to the agent. | Yes. All 50 IDs returned. | An agent cannot look up or report messages when it does not have their IDs. |
| Provider-hosted template send with variables | Yes | No native template-send or template-batch tool exposed | Postmark feature. It is not fair to put in a shared average. |
| Attachments | Not exposed by installed send tools | Yes | Resend feature. |
| Schedule, update, cancel email | No equivalent exposed | Yes | Resend lifecycle feature. |
| Explicit idempotency key | No explicit key exposed | Yes | Resend can explicitly support safe retry/deduplication patterns. |
| Default Reply-To configuration | No. Per-message Reply-To only. | Yes | Removed from shared tasks because this was not an even setup. |
| Inbound email, contacts, segments, broadcasts | No equivalent broad surface | Yes | Resend has a broader email-platform MCP surface. |
| Outbound diagnosis | Yes | Related logs/events/metrics tools | Not tested head to head because tool contracts differ. |
All direct latency benchmarks use 10-recipient batches because the Postmark MCP returned only 10 canonical IDs to the agent.
Both MCPs completed the shared GPT-5.6 task suite correctly. Postmark returned the initial acceptance response sooner, but its provider record was often not immediately readable, which led to more agent retries. Resend took slightly longer to return acceptance, but its record was almost always ready for the agent to inspect immediately. Postmark fits teams that send batches and use provider-hosted templates. Resend is a better fit for individual email workflows that need immediate record lookup, attachment support, or scheduling and cancellation.
You can build an eval like this one to check whether your agent sends the right content to the right recipients and verifies each email. Sign up for free to test your own workflow, or book a demo to walk through your setup.
A newsletter for unfiltered thoughts on eval methodology, analysis, and failures
Subscribe