Google’s multimodal AI assistant
Google’s multimodal AI for text, images and code, integrated across Google apps.
Understands and writes natural language — prompts in, prose out.
Reads images: photos, screenshots, diagrams and scanned documents.
Works with speech and sound as an input or an output, not just text.
Writes, explains and refactors source code across common languages.
Calls your own functions and tools with structured arguments.
Looks things up on the live web rather than answering from memory.
Guarantees valid, schema-shaped JSON your application can parse.
Streams the response token by token so output appears immediately.
ChatGPT
OpenAI Conversational AI assistant by OpenAI
4.8
Freemium
Compare
Claude
Anthropic Helpful, harmless AI assistant by Anthropic
4.8
Freemium
Compare
M
Midjourney
Midjourney Best-in-class AI image generation
4.7
Paid
Compare
E
ElevenLabs
ElevenLabs Realistic AI voice generation
4.7
Freemium
Compare
Llama
Meta Open foundation models by Meta
4.6
Free
Compare
ChatGPT
OpenAI
4.8
Claude
Anthropic
4.8
M
Midjourney
Midjourney
4.7
E
ElevenLabs
ElevenLabs
4.7
Llama
Meta
4.6
ChatGPT
OpenAI
4.8
2
Claude
Anthropic
4.8
3
M
Midjourney
Midjourney
4.7
4
E
ElevenLabs
ElevenLabs
4.7
5
Gemini
Google DeepMind
4.6
Head to the official site for docs and API keys, or compare Gemini against every other model in the directory.