Model profile

Muse Spark 1.3: Meta’s agent model, not the Muse app

Vendor table from Meta’s Muse Spark 1.3 card. Effort levels differ by column. Not an AIToolsRatings rerun. Model card.

Muse Spark 1.3 is Meta’s API model for long agent and coding runs. The Muse app that sends email and books travel is a different product, US-only, sitting on top of this model. This profile dates the vendor table and keeps the two names apart.

MetaDeveloper
Generative LLMShape
muse-spark-1.3Current id
$1.25 / $4.25Per 1M, standard*
1,048,576Context tokens*
Model card*

What Muse Spark is

Muse Spark 1.3 is Meta’s API model for long-horizon agent work and coding. It takes text, images, video, audio, and PDFs, and it returns text. The current id is muse-spark-1.3. Context on the model card is 1,048,576 tokens. It is a generator. A closed label, route, or allow/deny is still the wrong job for it — that page is Jev.

Meta also shipped a consumer product called Muse on 8 September 2026. That app messages you, opens a browser, and can send mail or pay, inside a vendor-described Secure VM. It is powered by Muse Spark. It is not this page. Mixing the names is how a $20-a-month consumer plan gets quoted as the API price. The split, including what Reuters reported about internal tests, is the desk note.

Developer and where you can call it

Meta. The builder surface is the Meta Model API, OpenAI-SDK compatible, described as a US developer preview. The model card also points at Muse Code, OpenRouter, and cookbooks for sandboxed computer use. The consumer surfaces — iOS, Android, muse.ai, and WhatsApp — are the app, currently rolling out in the United States, with AI glasses listed as later. Those are not API endpoints.

Two prices, and one of them sells your prompts

Standard muse-spark-1.3 is $1.25 input, $0.15 cached input, $4.25 output per million tokens. Meta’s card says this tier is not used to improve their products. The contributor id, muse-spark-1.3-contributor, is $0.10 / $0.002 / $0.20, and the card says that tier is used to improve their products. The cheap tier is a data trade, not a sale. Reasoning tokens are billed as output. The full arithmetic is on the pricing review.

How to read the vendor sheet

The table below is Meta’s Muse Spark 1.3 card. Effort is not matched: 1.3 is max, 1.2 is xhigh, GPT-5.6 Sol and Claude Opus 5 are max. One row, Agentic IF Index, is marked internal. Opus 5 has no MRCR cells on the card. We did not rerun any of it. Terminal-Bench here is 2.1. The Grok 4.7 profile quotes Terminal-Bench 4.0. Those are different tests. Do not subtract 88.8 from 38.0 and call it a gap.

Where 1.3 leads the card

  • Long context, MRCR. 98.5 in the 256K–512K band and 98.1 in the 512K–1M band. Sol is 91.5 and 73.8. Opus 5 is not given a number. This is the clearest cell in Spark’s favor, and it is still Meta’s harness.
  • DeepSWE v1.1. 75.4, ahead of Opus 5 at 74.0 and Sol at 73.0, and well ahead of Spark 1.2 at 55.0. A one-point edge over Opus is a near-tie. The jump from 1.2 is the number that matters if you already call Meta’s API.
  • SWEAtlas, questions about a codebase. 59.4 versus Sol 53.5 and Opus 52.7.
  • JobBench. 64.9 versus Sol 45.4. Opus 5 is still ahead at 65.7. Quote the Sol gap. Do not claim the Opus win.
  • OSWorld partial score versus 1.2. 66.9 versus 47.6. The binary score is 32.0. Computer-use “success” depends on which of those two numbers you print.

Where the same card is behind

  • GDPVal-AA v2. 1,754. Opus 5 is 1,824. Sol is 1,710, so Spark is between them and closer to Sol than to a trophy. Knowledge work is not the reason to leave Opus on this row.
  • OSWorld 2.0. Opus 5 is 68.3 partial and 31.4 binary. Spark is 66.9 and 32.0. Spark’s binary edge is small. Opus’s partial edge is also small. Neither is a migration by itself.
  • DeepSearchQA. Sol leads at 93.1. Spark is 90.3, Opus 90.4. Browsing agents are a near-tie on this card.
  • AutomationBench. 49.6, behind Opus at 50.3 and ahead of Sol at 46.7. End-to-end business workflows are not a Spark sweep.
  • Terminal-Bench 2.1. 88.8, tied with Sol, ahead of Opus at 86.7. A tie is a tie. It is also not the Terminal-Bench 4.0 number on the Grok page.
  • Agentic IF Index. 57.8, behind Sol at 60.5 and Opus at 59.1. Meta labels this one internal. An internal instruction-following index is the weakest cell to argue from.

What the consumer app claims, kept off this score

Meta says the Muse app runs in its own cloud VM, that a separate Sentinel must approve anything reaching the internet, that the model cannot see passwords or card numbers, and that it asks before sending mail or paying. It also says conversations in that VM are not shared with Meta’s ad systems, and that a Confidential VM encrypted to a key only the user holds comes later this year. Later this year is not shipped. Reuters reported on 8 September that internal tests had found the product stalling and exposing sensitive data without authorization. That is Reuters’ account of Meta’s tests, not a measurement by this desk. Safety marketing does not raise the model score above.

Input and output

  • Modalities in: text, image, video, audio, PDF. Out: text.
  • Reasoning effort is a request parameter. Deeper effort spends more output tokens, including reasoning tokens.
  • Tool calls and a browser inside a VM do not turn the model into a decision model.
  • The contributor checkpoint is the same family at a lower price because Meta may train on those prompts and completions.

When to pick something else

  • You wanted the personal agent that lives in WhatsApp and books travel. That is the Muse app, US-only at launch, and it is not an API swap. Read the desk note before you install it.
  • The artifact is a long document under Claude’s careful-reader default → Claude.
  • You are comparing coding agents on xAI’s September sheet, not this one → Grok 4.7.
  • You only needed a typed decision → Jev.

Latest update

23 September 2026 desk note. Muse Spark 1.3’s card and the 8 September consumer launch are both in view. This page rates the API model. It does not onboard a key, and it does not crown a leaderboard.

Task fit, not a winner

Benchmark (Meta card)Muse Spark 1.3 maxMuse Spark 1.2 xhighGPT-5.6 Sol maxClaude Opus 5 max
GDPVal-AA v2, knowledge work1,7541,6151,7101,824
JobBench, professional tool use64.961.645.465.7
OSWorld 2.0, partial / binary66.9 / 32.047.6 / 17.962.7 / 27.368.3 / 31.4
DeepSearchQA, agentic browsing90.385.993.190.4
Agentic IF Index (Meta internal)57.846.260.559.1
AutomationBench, end-to-end workflows49.638.246.750.3
MRCR 256K–512K98.566.391.5Not listed
MRCR 512K–1M98.155.573.8Not listed
DeepSWE v1.175.455.073.074.0
SWEAtlas, codebase questions59.446.253.552.7
Terminal-Bench 2.188.882.988.886.7
Get listed

Put your AI tool in front of people who are already comparing options.

Submit a listing for review. Complete submissions with a live website, pricing, and a clear use case typically go live within 24–72 hours.

Submit a tool