Meta's open-weight model family — the default foundation for self-hosted AI.
Conversational Chat
Multi-turn natural-language conversation with follow-up context.
Reasoning Mode
A deliberate "thinking" mode that works through multi-step problems before answering.
Multilingual Support
Understands and replies in dozens of languages, including translation between them.
Long-Context Understanding
Handles very long documents and conversations without losing earlier context.
Text Generation
Drafts, rewrites and edits long-form and short-form copy.
Image Understanding
Vision modelsReads photos, screenshots, diagrams and handwriting and answers questions about them.
Public API
Via providersA documented API for building the same capability into your own product.
Code Generation & Debugging
Writes, explains, reviews and fixes code across common languages.
Custom Assistants
Fine-tuningBuild and share a configured assistant with its own instructions and knowledge.
Task Automation
Schedules recurring jobs or runs multi-step tasks without supervision.
Third-Party Integrations
Huge ecosystemConnects to other SaaS products through apps, connectors or an integration platform.
Export & Download
Open weightsExport chats, documents or data out of the product in an open format.
Data-Retention Controls
Self-hostedConfigure how long conversation data is kept, or turn training off.
GDPR Compliance
Self-hostedOffers a data-processing agreement and EU data-protection commitments.
Encryption at Rest & in Transit
Traffic and stored data are encrypted with industry-standard ciphers.
Prices are set by the vendor and may change.
$0
What's included
Per token usage
What's included
Custom
What's included
No written reviews yet
Be the first to share your experience with Llama.
Llama is Meta's family of open-weight large language models. Meta publishes the weights under a permissive community licence, so anyone can download, run, fine-tune and deploy them without paying per token or sending data to a third party.
That decision reshaped the industry. Llama is now the default base model for self-hosted AI, and the ecosystem around it — Ollama, llama.cpp, vLLM, thousands of community fine-tunes — is larger than any other open family's.
Sizes range from small models that run on a laptop to large ones needing serious GPU clusters, so the same family covers edge devices and datacentre workloads.
Why it became the default:
Llama is infrastructure, not a consumer product.
If you just want to chat with an AI, use a hosted assistant — Llama requires engineering to become a product.
Longer context support across the model family.
Image understanding added to the open-weight releases.
Updated community licence terms and usage threshold.
Compact variants targeted at phones and edge hardware.
Join the users who trust Llama for their daily tasks.