1 October 2026

Resend vs Postmark MCP for transactional email

Jess Wang13 min
Key takeaways
240/240 emails accepted and verified with each MCP
Both passed all seven deterministic scorers across 90 tasks per provider in two runs. Acceptance and verification do not confirm inbox delivery.
0.16–0.22 seconds faster batch acceptance with Postmark vs Resend
Postmark returned acceptance confirmations faster on average when sending one, two, or five 10-email batches at a time.
99.6% first-lookup success with Resend vs 69.2% with Postmark
Resend needed a lookup retry for one of 240 emails. Postmark needed retries for 74, adding 98 lookup calls.
50 IDs returned by Resend vs 10 by Postmark
In the 50-email batch test, Resend returned every accepted email ID. Postmark exposed only the first 10 to the agent.

Transactional email delivers password resets, receipts, account alerts, and onboarding messages. At scale, teams rely on batches to send many personalized emails at once and templates to keep high-volume messages consistent. A pipeline failure can leave emails unsent or send an invoice or account detail to the wrong recipient.

AI agents and MCP servers can automate parts of that workflow. Transactional messages can contain password-reset links, invoice details, account information, and personal recipient data. This evaluation compares the Resend MCP and Postmark MCP across common tasks, direct acceptance timing, and record availability for follow-up lookup.

Methodology

  1. Full task evaluation: Codex completes nine email-task types through each MCP and returns a structured response. This measures workflow correctness and surfaces behavior in the traces.

  2. Direct benchmark: The task evaluation includes agent reasoning and tool-use time, so it cannot isolate MCP latency. The direct benchmark uses a non-agent client. It tests one 10-email batch at a time, then low- and high-concurrency bursts of two and five batches. Each measurement ends when the tool returns accepted IDs.

  3. Feature comparison: Some features could not be included in a fair head-to-head task because only one MCP exposed them. The feature comparison makes the scope of the two tools clear.

Source experiments and traces are available in the Braintrust project.

The email workflow

The following terms clarify the email pipeline used throughout the evaluation.

TermMeaning
Accepted IDThe provider's canonical identifier for one accepted email. Codex receives it after the provider accepts the send, uses that same ID for lookup and retries, and includes it in its final report so the harness can identify the exact email. Acceptance does not mean the email is in the recipient's inbox.
Provider recordThe provider's stored record for that accepted ID. It includes the recipient, subject, content, tag, and current provider state.
Status visibilityHow long after acceptance it takes for the same ID to return a usable record through the MCP lookup tool. Before then, the tool can return not found, unavailable, an error, or a response without a matching ID and usable state.
BatchOne native tool call containing separate messages for several recipients. Each recipient should receive a separate provider ID. A three-recipient batch should return three IDs.
Email unitOne required recipient-specific email. A 10-recipient batch is one task and one MCP call, but 10 email units.

Follow one email through the workflow

Codex receives an email task from the dataset and calls either the Resend MCP or Postmark MCP. The provider accepts the request and returns an email ID to Codex. Codex then calls the MCP again with that ID to retrieve the provider record. Finally, Codex returns structured JSON to the eval harness with the recipient, ID, and observed status.

Full task evaluation

Dataset

The dataset contains nine task buckets with five variants per bucket.

  • Password reset: Basic direct send and lookup.
  • Receipt: Tests whether a direct send preserves text, HTML, URL, visible link label, and tag.
  • Service notice: Tests that a direct send includes both required text and HTML, a precise maintenance window, a tag, and no unintended recipients.
  • Two-recipient welcome batch: Tests whether Ada and Lin's subjects, bodies, and URLs stay paired with the correct person in one batch call.
  • Two-recipient receipt batch: Tests that Sam and Rin receive their distinct invoice numbers, amounts, subjects, and body content in one batch call.
  • Content / serialization direct send: Tests whether Unicode, emoji, curly quotes, special characters, multiline text, and query-string links survive MCP serialization.
  • Shuffled five-recipient batch: Tests whether the MCP keeps each distinct subject, body, and HTML link paired with its intended recipient, even when the requested order is shuffled.
  • Shuffled ten-recipient batch: Tests the same mapping problem at 10 recipients.
  • Lookup-retry direct send: Tests that a temporary not found response causes Codex to retry only the lookup and not create a duplicate send.
The following examples show these tasks and their potential failure modes.
Task shapeOne controlled exampleWhat a failure would look like
Content / serialization direct sendSend Ada ExampleCo café update ☕ — 01. The body starts Hi Zoë 👋, and includes https://example.test/reports/cafe?locale=fr&source=v2. The HTML link is labelled Open report. Other variants use curly quotes, &, <v2.1>, Japanese text, apostrophes, and multiline copy.café becomes café. Emoji or curly quotes disappear. <v2.1> is interpreted as markup rather than text, a blank line disappears, or the link loses &source=v2.
Shuffled five-recipient batchThe requested order is Sam, Lin, Ada, Mika, then Rin. Sam receives ExampleCo mapping MAP-5-01-03 for Sam, amount $34.25, and a Review MAP-5-01-03 link with recipient=3. Every other recipient has different copy and a different link.All five emails are sent, but Sam receives Lin's amount or link. Two payloads are swapped, one recipient is missing, or the agent sends five individual messages instead of one native batch.
Two-recipient welcome batchIn one batch, Ada receives Welcome to ExampleCo, Ada and https://example.test/dashboard/ada?cohort=1. Lin receives the Lin subject and https://example.test/dashboard/lin?cohort=1.Lin gets Ada's dashboard URL or name. Both recipients receive identical content, the agent combines the recipients in one email, or it uses two ordinary sends rather than one native batch.

Scorers

The evaluation uses seven deterministic scorers to test different parts of this pipeline.

ScorerQuestion
provider_send_acceptedDid the provider accept every required email and return an ID?
agent_status_retrievedDid Codex itself get a usable record for every ID before it ended? A usable record has the matching provider ID and a real state, such as Sent or delivered.
direct_payload_fidelityDid Codex pass the requested subject, text, HTML/link, and tag to the MCP?
provider_payload_fidelityDid the provider record retain the required sender, content, and tag?
batch_semanticsDid Codex use one native batch with distinct messages, rather than a sequence of sends?
recipient_safetyDid every message go only to the intended recipient, with no CC/BCC or cross-recipient mix-up?
final_report_identityDid the final report include one correct recipient and canonical provider ID for every expected email?

This diagram shows how each deterministic scorer verifies a specific behavior in the email workflow.

Chronological pipeline showing where each deterministic scorer runsCodex sends an email through the MCP, the provider returns an accepted ID, Codex looks up the same ID, reports it to the eval harness, and the harness fact-checks the stored record through the provider REST API. Seven scorers check the send payload, batching, acceptance, lookup, final report, stored payload, and recipients.Codex agentEmail MCPPostmark MCP or Resend MCPEmail providerPostmark or ResendEval harnessTask: send the requested email, then check it and report the IDdirect_payload_fidelity: requested subject, text, HTML/link, and tagbatch_semantics: native batch operation, not sequential sendsprovider_send_accepted:every requested email was accepted and has a canonical IDagent_status_retrieved: Codex received a usable same-ID provider recordfinal_report_identity: one correct recipient + canonical ID per emailprovider_payload_fidelity: stored sender, content, and tagrecipient_safety: no unexpected recipient, CC/BCC, or cross-exposureSend email requestProvider API send requestAccepted + canonical email IDReturn accepted IDLook up the same IDProvider-record lookupRecord or temporary-unavailable responseReturn lookup resultFinal JSON: recipient + canonical IDDirect REST API fact-checkCanonical stored record

Results

Both MCPs passed every check

The task evaluation ran twice with the same dataset and configuration. The results below combine both runs.

ScorerPostmarkResend
provider_send_accepted240/240240/240
agent_status_retrieved240/240240/240
batch_semantics40/40 batch task rows40/40 batch task rows
final_report_identity90/9090/90
direct_payload_fidelity950/950 requested send-field comparisons950/950 requested send-field comparisons
provider_payload_fidelity1,190/1,190 requested stored-field comparisons1,190/1,190 requested stored-field comparisons
recipient_safety90/90 task rows90/90 task rows

These deterministic checks show that both MCPs completed the tasks correctly. They do not identify a clear winner between Resend and Postmark.

Resend returned usable records sooner

First-lookup success differed between the two providers. It measures whether the first MCP lookup after the provider accepted an email returned a usable provider record rather than a temporary not found response.

MeasurePostmarkResend
First same-ID lookup usable166/240 = 69.2%239/240 = 99.6%

The difference in first-lookup success was 30.4 percentage points. Postmark often accepted the email before the agent could retrieve its record to verify content and status. This matters for two reasons. The agent must retry the lookup, which adds latency and cost. An agent that interprets the missing record as a failed send could incorrectly report failure or send a duplicate email.

Postmark required more lookup retries

For Postmark, the first lookup for 74/240 email units returned an unavailable result and needed a retry, which led to 98 extra Postmark lookup calls.

If we group email units by the number in each send-tool call, larger batches had fewer Postmark first-lookup misses, probably because they give records more time to become available before Codex looks them up. In one 10-recipient Postmark trace, the batch was accepted at 52.12s and the first lookup began 15.04s later, giving the provider record time to appear.

Email units in the original send callPostmark email units needing a retryPostmark median time to first successful lookupResend email units needing a retry
142/50 (84%)9.68s0/50
220/40 (50%)12.17s0/40
512/50 (24%)10.81s0/50
100/100 (0%)15.04s1/100

Checking the delay without the MCP

The next question was whether the delay came from the MCP wrapper or the provider. A direct REST control called each provider's API with no Codex agent or MCP server. It used one task from each of the nine task buckets.

MeasurePostmarkResend
Accepted units24/2424/24
Median REST status-visibility time11.79s1.27s

The delay before the first readable provider record remained even after the MCP and agent were removed. This indicates that the observed delay comes from the Postmark provider rather than the Postmark MCP.

Estimated model cost

To calculate the estimated cost, we used the traces' token counts. The estimate uses the current GPT-5.6 Sol prices of $4 per million input tokens, $0.40 per million cached input tokens, and $20 per million output tokens (GPT-5.6 Sol pricing). It excludes Postmark and Resend plan charges.

MeasurePostmark total / avgResend total / avg
Prompt tokens10.64M9.55M
Cached prompt tokens9.43M8.45M
Output tokens92.8K100.1K
Reconstructed LLM cost$10.45$9.80
Reconstructed LLM cost per task$0.116$0.109

Based on the rounded totals above, Postmark's estimated LLM cost was about $0.65 higher across the entire comparison, or about $0.007 more per task. This is probably because Postmark made 68 more model calls, consistent with extra lookup retries caused by the delay between email acceptance and provider-record availability.

Direct benchmark

The second part of this head-to-head comparison calls each MCP with no agent involved, which isolates MCP behavior.

The benchmark used three test conditions.

  • Single batch: One 10-email batch call.
  • Low concurrency: Two 10-email batch calls at the same time, for 20 email units total.
  • High concurrency: Five 10-email batch calls at the same time, for 50 email units total.

The single-batch test is the baseline latency condition. It measures one batch-acceptance round trip without competing calls. The low- and high-concurrency tests measure burst behavior, testing whether the MCP and provider stay responsive when two or five clients submit batches at the same time.

Postmark accepted batches sooner

Test conditionResult
Single batchPostmark returned accepted IDs faster in all 30 matched 10-email batch calls, averaging 0.156 seconds sooner (95% CI, 0.095–0.218 seconds).
Low concurrencyPostmark was faster in all 25 matched waves, averaging 0.219 seconds sooner (95% CI, 0.072–0.366 seconds).
High concurrencyPostmark was faster in all 25 matched burst waves, averaging 0.173 seconds sooner (95% CI, 0.154–0.192 seconds).

Postmark returned accepted IDs sooner after a direct batch request. Resend made those accepted email records available sooner when Codex immediately tried to look them up, which reduced retry work.

Direct MCP batch acceptance time
Lower is faster. Shaded bands show 95% confidence intervals across three sessions.
Acceptance round-trip time
0.0s0.1s0.2s0.3s0.4s0.5s0.6s0.7s
Single batch
30 matched calls
Low concurrency
25 matched waves
High concurrency
25 matched waves
Postmark
Resend

Feature comparison

Some features could not be evaluated head to head because only one MCP exposed them. The differences were substantial enough to document separately.

FeaturePostmark MCPResend MCPWhy it matters
Direct transactional sendYesYesShared workflow tested head to head.
Native direct-content batchYes, up to 500 provider itemsYes, up to 100 itemsShared workflow tested up to 50 items.
Return all IDs from a 50-item successful batchNo. Only the first 10 IDs were shown to the agent.Yes. All 50 IDs returned.An agent cannot look up or report messages when it does not have their IDs.
Provider-hosted template send with variablesYesNo native template-send or template-batch tool exposedPostmark feature. It is not fair to put in a shared average.
AttachmentsNot exposed by installed send toolsYesResend feature.
Schedule, update, cancel emailNo equivalent exposedYesResend lifecycle feature.
Explicit idempotency keyNo explicit key exposedYesResend can explicitly support safe retry/deduplication patterns.
Default Reply-To configurationNo. Per-message Reply-To only.YesRemoved from shared tasks because this was not an even setup.
Inbound email, contacts, segments, broadcastsNo equivalent broad surfaceYesResend has a broader email-platform MCP surface.
Outbound diagnosisYesRelated logs/events/metrics toolsNot tested head to head because tool contracts differ.

All direct latency benchmarks use 10-recipient batches because the Postmark MCP returned only 10 canonical IDs to the agent.

Conclusion

Both MCPs completed the shared GPT-5.6 task suite correctly. Postmark returned the initial acceptance response sooner, but its provider record was often not immediately readable, which led to more agent retries. Resend took slightly longer to return acceptance, but its record was almost always ready for the agent to inspect immediately. Postmark fits teams that send batches and use provider-hosted templates. Resend is a better fit for individual email workflows that need immediate record lookup, attachment support, or scheduling and cancellation.


You can build an eval like this one to check whether your agent sends the right content to the right recipients and verifies each email. Sign up for free to test your own workflow, or book a demo to walk through your setup.

Share

Read more evals

GPT-6.1 Sol vs GPT-6 Sol for problem solving and teaching
30 September 2026
Opus 5.5 vs new GPT-6 models for writing quality
22 September 2026
When Jev holds up as a judge
21 September 2026

Subscribe to the Department of Evals

A newsletter for unfiltered thoughts on eval methodology, analysis, and failures

Subscribe