SIMA 2

SIMA 2

Google DeepMind · Image

SIMA 2 is Google DeepMind's AI agent that plays, reasons, and learns alongside you inside 3D virtual worlds. Unlike a simple command follower, it runs on Gemini models to understand your goals, explain its plans, and chat about what it's doing, all while driving a character with the same keyboard and mouse inputs a person would use. It's a research preview, not a product you can download today. Think of it as a virtual world AI agent built for study, not a tool waiting in an app store.

Interface preview of SIMA 2

About SIMA 2

What Is SIMA 2

SIMA 2 stands for Scalable Instructable Multiworld Agent. It's the follow-up to the original SIMA from 2024, and it marks a shift from an AI that just follows orders to something closer to a collaborative game partner. The earlier version handled basic commands across a handful of commercial games. SIMA 2 adds reasoning, dialogue, and the ability to pick up tricks in places it has never seen before.

The core change is Gemini. Where SIMA 1 learned a direct link between words and button presses, SIMA 2 embeds a Gemini model as its reasoning engine. That lets it break a high-level goal into steps, narrate its intent, and correct course when something goes wrong. Language becomes intent. Intent becomes a plan. Then a plan becomes action.

There's a real limitation worth stating up front. SIMA 2 is a research preview from Google DeepMind, not a consumer app. There's no download, no sign-up page, and no public API. You can't try it on your own machine yet. The team says it's meant to study how agents transfer skills from virtual environments toward real robots, and it deliberately stays at the high-level planning layer rather than controlling physical joints or wheels.

Getting Started

  1. Understand there's no public product. Track progress through the Google DeepMind research page and official announcements.
  2. Read the published research on how the agent pairs a Gemini model with an embodied skill set.
  3. Watch the demo footage showing the difference between SIMA 1 and SIMA 2 handling open-ended requests in games like No Man's Sky.
  4. Follow the team's updates on Genie and robotics work if you want to see where the research heads next.

Product Information

A quick look at SIMA 2's pricing, supported platforms, and performance.

Free PlanNo
Paid PlansNot available to the public
PlatformWeb (research preview, no public download)
DeveloperGoogle DeepMind
CategoryImage
Release DateNov 2025
Latest UpdatedNov 2025
Website VisitsN/A
Website Global RankN/A
API AvailabilityNo

Best for

The users, tasks, and scenarios where this tool fits best.

Users

  • AI researchers
  • Robotics teams
  • Game AI enthusiasts

Tasks

  • Studying multi-step planning
  • Testing generalization
  • Exploring self-improvement

Scenarios

  • Academic literature review
  • Product research
  • Demonstrations and talks

Key features

Gemini as the Reasoning Engine

SIMA 2 puts a Gemini model at the center of how it thinks. This AI agent for 3D worlds uses Gemini to interpret a goal, work out a plan, and talk through what it's doing. The published work points to the Gemini 2.5 Flash-Lite version as the engine inside the agent.

Natural Language Dialogue

You can talk to SIMA 2 the way you'd talk to a teammate. It answers in plain language, describes its intent, and lays out the steps it plans to take. No scripted barks. No canned lines. The intent shifts from handing out orders to working on a problem together, which is the clearest break from the earlier SIMA.

Multimodal Instructions

The agent reads more than typed words. It responds to voice, hand-drawn sketches, and even emoji, so you can point to a spot on the map or send a pair of icons to signal an action. Once you get past the novelty, the sketch input is the practical one: it's a fast way to mark a location or a path without writing a sentence.

Skill Transfer Across Games

SIMA 2 carries concepts from one world into another. Learn how to mine in one game, and it can apply that idea to gathering in a different one. That transfer is what makes the agent useful research: skills stop being tied to a single title.

Self-Improvement Loop

After the initial human demonstrations, the system generates its own tasks and tests them, leaning on Gemini feedback to sharpen its approach. It's early, and the gains are measured in a research setting, but the loop is the part that separates it from a static model.

Higher Task Success

According to Google DeepMind, SIMA 2 reached roughly 65% success on games it trained on, close to a human baseline near 75%. The first SIMA managed about 31% on complex tasks, so the jump is large. Treat these as the lab's own numbers, not an independent verdict.

Works From Pixels, Not Code

The agent doesn't touch a game's source code or special API. It watches the screen and sends virtual keyboard and mouse signals, the same way a person does. That constraint is why the approach might one day extend beyond games.

Pros and cons

Pros

  • Reasons about goals instead of only following literal commands, so it handles looser instructions.
  • Explains its plan in words, which makes its behavior easier to follow and trust.
  • Reads text, voice, sketches, and emoji, giving you several ways to give it direction.
  • Transfers learned skills to new games, a real step past single-title agents.
  • Runs without game code or custom APIs, so it can work with existing titles.

Cons

  • It's a research preview only. There's no download, no sign-up, and no public API, so you can't use it yourself.
  • Success rates come from Google DeepMind, and no independent benchmark backs them up yet.
  • It handles high-level decisions but not fine physical control, so it's far from a finished robotics system.
  • No confirmed release date exists, so any plan built around it amounts to guesswork.

Frequently asked questions

No. It's a research preview from Google DeepMind, with no consumer app, no download link, and no public API. Published research and official demos are the only way to follow along.