OpenAI's text-to-video model with synchronised dialogue and sound. The Sora apps closed on 26 April 2026 and the sora-2 API retires on 24 September 2026.
Sora is OpenAI's video generation model. It turns a written prompt, or a single reference image, into a short photorealistic clip with camera movement, physically plausible motion and, from Sora 2 onwards, a synchronised soundtrack of dialogue, ambience and effects. It is the model that put AI video in front of a mainstream audience.
OpenAI previewed Sora on 15 February 2024 and opened it to ChatGPT Plus and Pro subscribers on 9 December 2024 as Sora Turbo, generating up to 20 seconds at 1080p for Pro accounts. Sora 2 followed on 30 September 2025 with a standalone social app, markedly better physical realism, native audio and cameos, which dropped a verified likeness into a generated scene. The Videos API arrived a week later and exposed sora-2 and sora-2-pro to developers.
Sora is now being retired. OpenAI announced the wind-down on 24 March 2026, closed the Sora web and mobile apps on 26 April 2026, and will remove the Videos API, including sora-2, sora-2-pro and every dated snapshot, on 24 September 2026. The company said the Sora research team is refocusing on world-simulation work to advance robotics, and that compute had to be traded against products with lower cost per unit of value. No direct replacement has been announced.
Until the shutdown, sora-2 and sora-2-pro still generate 720p to 1080p clips from text or from an image used as the first frame, up to 20 seconds per call, with extensions chaining as many as six continuations into a two-minute sequence. The guardrails are strict: real people including public figures cannot be generated, copyrighted characters and music are rejected, input images containing faces are refused, and output is limited to content suitable for audiences under 18.
If you are building on Sora today, plan the migration now. Runway, Pika, Google's Veo, Kling and Luma's Dream Machine cover the same text-to-video ground, and several of them match the synchronised audio and image conditioning that made Sora 2 distinctive. The endpoint disappears on 24 September 2026 rather than degrading gracefully.
Text to video
A written prompt becomes a shot: framing, camera move, lighting and motion in one pass.
Native synchronised audio
Sora 2 generates dialogue, ambience and effects locked to the picture, not dubbed after.
Image conditioning
Supply a JPEG, PNG or WebP at the target resolution and Sora uses it as the first frame.
20s, extendable to 2 min
Twenty seconds per call, chained through up to six extensions to 120 seconds total.
Understands and writes natural language — prompts in, prose out.
Reads images: photos, screenshots, diagrams and scanned documents.
Works with speech and sound as an input or an output, not just text.
Text-to-video generation
Describe a scene in natural language and Sora returns a finished clip: subject, camera movement, lighting and motion in a single pass, with no keyframing and no editing timeline.
Synchronised dialogue and sound
Sora 2 was the first OpenAI model to generate a soundtrack with the picture. Spoken lines, room tone, footsteps and effects are timed to the action rather than added afterwards.
Image-to-video
Pass a JPEG, PNG or WebP that matches the target video resolution and the model treats it as the first frame, carrying its composition, palette and subject into the motion.
Extensions up to two minutes
A completed clip can be continued by up to 20 seconds at a time, to a maximum of six extensions or 120 seconds. Extensions do not carry characters or image references forward.
Physical plausibility
Sora 2 holds objects, gravity and contact consistent across a shot, cutting the melting-and-morphing artefacts that gave earlier video models away within a second or two.
Guardrails on every generation
Real people including public figures cannot be generated, copyrighted characters and music are rejected, input images with faces are refused, and output is limited to material suitable for audiences under 18.
Sora is billed per second of finished video, not per token. Standard API rates are below; batch runs at half price. Access ends when the Videos API is removed on 24 September 2026.
| Provider | Input | Output | Cached Input | Batch | Fine-Tuning |
|---|---|---|---|---|---|
| sora-2 — 720p | — | $0.10 / sec | — | $0.05 / sec | Not available |
| sora-2-pro — 720p | — | $0.30 / sec | — | $0.15 / sec | Not available |
| sora-2-pro — 1024p | — | $0.50 / sec | — | $0.25 / sec | Not available |
| sora-2-pro — 1080p | — | $0.70 / sec | — | $0.35 / sec | Not available |
Head to the official site for docs and API keys, or compare Sora against every other model in the directory.