llm-bill

Independent test · measured 2026-10-10 to 2026-10-11 · 日本語版

Claude Haiku 5.5 vs GPT-6 Luna

Both list at $0.10 per million input tokens and $0.50 per million output tokens. We ran 50 support tickets, 50 invoices and 25 email requests through both and counted every billed token. Haiku answered faster on every task we ran. Luna was cheaper on 2 of 3 tasks, and both blind judges preferred its Japanese business emails on our 25 requests.

Which one to use

Based only on what we measured. If your workload looks different, put your own numbers into the cost calculator.

JobPickWhy
Short classification or routing where latency mattersClaude Haiku 5.5Cost within a few percent (×0.98). Median 0.90 s vs 1.45 s, p90 1.08 s vs 2.08 s.
Field extraction with a JSON schemaGPT-6 LunaSame accuracy, 39% cheaper. The extraction schema alone adds 348 input tokens per request on Haiku against 68 on Luna.
Japanese business email and other customer-facing Japanese textGPT-6 LunaPreferred by both blind judges and 33% cheaper. Haiku wrote 49% longer emails.
Prompts longer than about 259,577 English charactersGPT-6 LunaHaiku's higher price tier starts at about 259,577 English characters per request; Luna's at about 1,133,815.
You are already on Claude and need the speedClaude Haiku 5.5Put today's date in the system prompt, and budget ×1.26–×1.62 the tokens an OpenAI-based estimate gives you.

The tokenizer gap

The two models list the same price per token, so the bill depends on how many tokens each one counts for the same text. We sent the same samples to both vendors' official token-counting endpoints: 24 each of English, Japanese and Chinese business text across 8 genres, written for this test, plus excerpts of real open-source files (24 Python, 24 JavaScript/TypeScript, 15 JSON). Sources and licenses

Claude Haiku 5.5 tokens per GPT-6 Luna token, same textDot is the pooled ratio; line is the 95% bootstrap interval. Hover or focus a dot for the numbers.
equal countsJavaScript / TypeScript×1.62JSON×1.61English×1.60Python×1.56Chinese×1.39Japanese×1.26
Data table
ContentsamplesRatio95% interval
JavaScript / TypeScript24×1.6161.559–1.681
JSON15×1.6081.517–1.690
English24×1.6041.552–1.683
Python24×1.5611.527–1.592
Chinese24×1.3931.375–1.413
Japanese24×1.2611.235–1.286

Japanese has the smallest gap, and it varies by genre: from ×1.15 (news article) to ×1.34 (support chat), 3 samples per genre. The worst case we saw was an all-caps legal clause (a single sample): ×2.36.

The gap applies to output too. When Haiku writes the same English reply, it is billed ×1.60 the output tokens, and output lists at $0.50 per million against $0.10 for input.

JSON schemas cost more on Haiku

Both vendors bill the instructions they add for structured output as input tokens. We measured how many by counting the same request with and without schemas of string fields.

SchemaClaude Haiku 5.5GPT-6 Luna
1 string field+188+23
6 string fields+318+68
20 string fields+682+194
Our 6-field extraction schema+348+68

Input tokens added per request. Roughly 162 + 26 per field on Haiku and 14 + 9 per field on Luna; enums and nesting add more.

Three real tasks

Same prompts, both models at medium effort, run 2026-10-10 to 2026-10-11. Costs are recomputed from the tokens each API billed, including reasoning tokens, at current list prices. Prompts

Cost per 1,000 requests (USD)Billed input, output and reasoning tokens at list price.
Ticket classificationClaude Haiku 5.5$0.033GPT-6 Luna$0.033Invoice extractionClaude Haiku 5.5$0.147GPT-6 Luna$0.090Japanese business emailClaude Haiku 5.5$0.237GPT-6 Luna$0.159
Data table
TaskClaude Haiku 5.5GPT-6 Luna
Ticket classification$0.033$0.033
Invoice extraction$0.147$0.090
Japanese business email$0.237$0.159
TaskSamplesClaude Haiku 5.5 costGPT-6 Luna costAccuracyMedian latencyp90 latency
Ticket classification — 5 labels; 25 English, 25 Japanese50$0.033$0.033100% / 100%0.90 s / 1.45 s1.08 s / 2.08 s
Invoice extraction — JSON schema, 6 fields; hard cases: 8 Japanese, 8 English50$0.147$0.090100% / 100%1.72 s / 2.23 s2.91 s / 2.86 s
Japanese business email — 25 scenarios25$0.237$0.159see below2.44 s / 3.56 s3.10 s / 4.55 s

Cost per 1,000 requests. Accuracy and latency are listed Haiku / Luna.

Classification comes out close because two effects offset each other. Haiku counted 48% more input tokens, while Luna spent about 17.9 reasoning tokens per ticket against Haiku's 3.0, and reasoning bills at the output price. By language, Luna cost $0.030 vs Haiku's $0.033 on the English tickets and $0.036 vs $0.033 on the Japanese ones, where it used 23.3 reasoning tokens per ticket against 12.5 in English.

Reasoning on extraction depended on the language too: Haiku spent about 69 reasoning tokens per Japanese invoice against Luna's 37, but 16 per English invoice against Luna's 28. assumption Haiku's figure is an estimate: billed output tokens minus the counted visible answer.

The hard extraction cases included totals given only as subtotal plus tax, amounts corrected later in the same thread, Japanese era dates and "net 30" due dates. Both models got all of them right, so this set does not separate them on accuracy.

Japanese business email

Each model drafted 25 emails from short requests of the kind people type into an assistant: apologies, quotes, payment reminders, holiday notices. Two judges from different vendors, Claude Opus 5.5 and GPT-6.1 Sol, saw each pair as A and B without model names, once in each order. A model counts as preferred only if the two orders do not cancel out. Judging protocol

Which email would you rather send?25 pairs per judge, after combining both A/B orders.
Claude Opus 5.5Luna 22Tie 2Haiku 1GPT-6.1 SolLuna 25
Data table
JudgeLunaTieHaikun
Claude Opus 5.5222125
GPT-6.1 Sol250025
Scores and agreement
JudgeLuna / tie / HaikuSame verdict in both orders"Send as-is", Haiku"Send as-is", Luna
Claude Opus 5.522 / 2 / 123 of 253.244.60
GPT-6.1 Sol25 / 0 / 025 of 253.344.74

Scores are 1–5 averages over both orders. The two judges agreed on 22 of 25 pairs.

The Anthropic judge preferred the OpenAI model. Three patterns stand out in Haiku's emails:

  • Haiku's emails were 49% longer, 442 characters against 297 on average. The judges listed specifics or commitments the request never asked for 60 times in Haiku's emails against 4 in Luna's (Claude Opus 5.5), and 55 times in Haiku's emails against 2 in Luna's (GPT-6.1 Sol), summed over both orders.
  • It left fill-in placeholders such as 「(連絡先:)」 or 「株式会社○○」 in 9 of 25 emails, against 3 for Luna. Both models left one in mail-09, mail-15, mail-18; two of these requests (mail-15, mail-18) asked for a link without giving the URL. How placeholders are counted
  • In 7 of 25 emails Haiku wrote a weekday or year inconsistent with the 2026 and 2027 calendars; Luna did in 0. Both counts come from a script that checks every date written with a weekday or year.

mail-21 · Year-end holiday notice

The request (Japanese)

取引先各位あての年末年始休業のお知らせ。休業期間は12/27(日)〜1/4(月)で、1/5(火)から通常営業。12/25(金)15時までの受注分は年内に出荷、それ以降は1/6以降の出荷になる。差出人は株式会社ミヤビ化成 業務部。

Claude Haiku 5.5

株式会社ミヤビ化成 各位
■ 休業期間:2024年12月27日(金)〜2025年1月4日(土)
 ※ご指定の期間は12/27(日)〜1/4(月)です
■ 通常営業再開:2025年1月5日(日)

GPT-6 Luna

取引先各位
12月27日(日)~1月4日(月)
1月5日(火)より通常営業いたします。

Problems marked above

  • Addressed to the sender's own company instead of its clients (取引先各位).
  • Moved the user's dates to 2024–25 years and weekdays; the request's 12/27 (Sun) to 1/4 (Mon) matches the 2026 calendar.
  • Left an internal note that contradicts the dates above it.

The date errors have a simple fix. We reran the 7 affected requests with 「今日は2026年10月10日(土)です。」 added to the system prompt, 2 times each, and every date and weekday came out right (14 of 14 outputs, checked by the same script). Luna made no date errors without it. If you use Haiku for anything with dates, give it today's date.

Long prompts: the price tier

Claude Haiku 5.5 bills the whole request at $0.50 / $2.50 per million tokens (instead of $0.10 / $0.50) once a single request's input passes 100K tokens. GPT-6 Luna moves the whole request to $0.20 / $0.75 above 272K. Because Haiku also counts more tokens for the same text, the gap in characters is wider than those numbers suggest. For English with a 400-character system prompt, Haiku crosses its threshold at about 259,577 characters per request and Luna at about 1,133,815.

As an example, summarizing a 300,000-character English document into Japanese costs ×7.87 as much on Haiku: about $30.64 vs $3.89 per 1,000 requests with Batch pricing. assumption Output length and reasoning tokens are assumptions, not measurements. The reasoning numbers reuse the email task's averages.

Try your own input size, languages, schema and cache rate.Open the cost calculator

Method and limits

  1. Models: claude-haiku-5-5 and gpt-6-luna through their official APIs, both at medium effort, with identical system prompts. Neither prompt included the current date. Models · Prompts
  2. Data: the 50 tickets, 50 invoices, 25 email requests and the English, Japanese and Chinese tokenizer samples were written for this test by AI agents. A script checks every dataset (labels, extraction spans, ids), and all checks pass. Code and JSON samples are excerpts of real open-source files, kept with their licenses. Datasets · Labels
  3. Prices: list prices from the vendors' official pricing pages, checked 2026-10-11. assumption Claude Haiku 5.5: The batch table lists only input and output prices; we assume the 50% batch discount also applies to cache reads ("Batch API requests are 50% off"). Cache-write costs are not included. Prices and sources
  4. Judging: the email preference comes from AI judges, not native-speaker reviewers. Both judges saw anonymized pairs in both orders. The date and placeholder counts come from scripts, not hand checks. Judging
  5. Size: 50 + 50 + 25 requests per model, 1 run each. Latency is from a single client location and will vary with your network and time of day. Limits
  6. Cost of this test: about $1.30 in API fees, including judging. Cost · Raw data · Rerun it
  7. Numbers: every number on this page is checked against the data files when the site is built. How