Ship a Full-Stack App with One Prompt

Copy this prompt into your AI coding agent, or open it in one below.

Give this to your AI Create a to-do list app using Puter.js

Coding manually? see the guide

Meta Llama

Meta Llama API

Access Meta Llama instantly with Puter.js, and add AI to any app in a few lines of code without backend or API keys.

// npm install @heyputer/puter.js
import { puter } from '@heyputer/puter.js';

puter.ai.chat("Explain AI like I'm five!", {
    model: "meta-llama/llama-4-maverick"
}).then(response => {
    console.log(response);
});
<html>
<body>
    <script src="https://js.puter.com/v2/"></script>
    <script>
        puter.ai.chat("Explain AI like I'm five!", {
            model: "meta-llama/llama-4-maverick"
        }).then(response => {
            console.log(response);
        });
    </script>
</body>
</html>

List of Meta Llama Models

Chat

Muse Glimmer 30B

meta/muse-glimmer-30b

Muse Glimmer 30B is a 30-billion-parameter dense model from Meta Superintelligence Labs, pairing a causal transformer with a 1.8-billion-parameter vision encoder for text and image input. It's released under an Apache 2.0 license, Meta's first fully open-weight model since Muse Spark moved to a paid API. On Meta's own benchmarks, it scores 76.0 on SWE-Bench Verified, 51.2 on SWE-Bench Pro, 94.7 on AIME 2026, 83.5 on GPQA Diamond, 75.5 on MCP Atlas, and 74.6 on DeepSearch QA, ahead of similarly sized open models like Gemma4 31B and Qwen3.6 27B on Meta's reporting. These figures are vendor-reported and not independently verified. Where Muse Spark targets large-scale multi-agent orchestration, Muse Glimmer sits a size tier down, aimed at tool use, multi-step reasoning, coding, and LLM-as-a-judge evaluation with a 131,072-token context window. It fits developers who want agentic and coding capability at lower cost than the larger Muse Spark models.

Chat

Muse Spark 1.2

meta/muse-spark-1.2

Muse Spark 1.2 is Meta Superintelligence Labs' coding-focused update to Muse Spark 1.1, released alongside Muse Code, a terminal coding agent it powers. Meta scaled up training compute on coding tasks and training-environment diversity, aiming at code generation, debugging, codebase understanding, and long-horizon work like whole-repository generation. On Meta's own evaluation harness, Muse Spark 1.2 scored 82.9% on Terminal-Bench 2.1 and 59.3% on DeepSWE 1.1, edging OpenAI's GPT-5.6 Terra (81.8%) and xAI's Grok 4.5 (81.6%) on Terminal-Bench but trailing Anthropic's Opus 5 (86.7%). These are vendor-run numbers, not yet independently verified on the public leaderboards. It keeps the 1,048,576-token context window and text, image, video, audio, and PDF input from 1.1, at the same $1.25 per million input and $4.25 per million output token pricing. A separate contributor tier offers lower rates in exchange for letting Meta train on your prompts and completions.

Chat

Muse Spark 1.1

meta/muse-spark-1.1

Muse Spark 1.1 is a multimodal reasoning model from Meta Superintelligence Labs, built for agentic workflows. It accepts text, images, video, audio, and PDF documents as input and returns text, with a 1,048,576-token context window. The model is designed to orchestrate multi-agent workflows, acting as either a main agent that plans and delegates tasks or as a subagent, and it generalizes zero-shot to new tools, MCP servers, and custom skills. It supports parallel function calling, structured output, built-in search with citations, and configurable reasoning effort, and Meta reports strong results on coding across large codebases, computer-use tasks, and visual-to-code generation. This is Meta's first model available through a paid API, priced at $1.25 per million input tokens and $4.25 per million output tokens, aimed at developers building agentic coding tools and enterprise workflow automation.

Chat

Llama Guard 4 12B

meta-llama/llama-guard-4-12b

Llama Guard 4 12B is Meta's 12 billion parameter multimodal safety model that moderates both text and image inputs across 12 languages. It was built from Llama 4 Scout and detects violations based on the MLCommons hazard taxonomy.

Chat

Llama 4 Maverick

meta-llama/llama-4-maverick

Llama 4 Maverick is Meta's 400 billion total parameter MoE model with 17B active parameters and 128 experts, supporting 1M token context. It's natively multimodal with state-of-the-art performance on coding, reasoning, and image understanding tasks.

Chat

Llama 4 Scout

meta-llama/llama-4-scout

Llama 4 Scout is Meta's efficient 109 billion parameter MoE model with 17B active parameters and 16 experts, featuring an industry-leading 10M token context window. It fits on a single H100 GPU and handles multimodal text and image inputs.

Chat

Llama Guard 3 8B

meta-llama/llama-guard-3-8b

Llama Guard 3 8B is Meta's enhanced safety moderation model providing content classification in 8 languages with support for tool call safety. It detects 14 hazard categories and integrates with Llama 3.1 for comprehensive AI safety.

Chat

Llama 3.3 70B Instruct

meta-llama/llama-3.3-70b-instruct

Llama 3.3 70B Instruct is Meta's refined 70 billion parameter multilingual model with improved instruction following and tool use capabilities. It supports 8 languages and offers enhanced reasoning performance over previous versions.

Chat

Meta Llama 3.3 70B Instruct Turbo

meta-llama/llama-3.3-70b-instruct-turbo

Llama 3.3 70B Instruct Turbo is a 70-billion-parameter, instruction-tuned text model from Meta, served on Together AI's throughput-optimized Turbo endpoint. It is known for delivering quality close to the much larger Llama 3.1 405B at a fraction of the cost, especially on instruction following and reasoning. Published benchmarks include IFEval 92.1%, MMLU 86.0%, HumanEval 88.4%, MGSM 91.1%, MATH 77.0%, and GPQA Diamond 50.5%. On IFEval it edges out Llama 3.1 405B (88.6) and approaches Claude 3.5 Sonnet (89.3). With a 128K context window and tool/function calling support, it suits developers building chat assistants, multilingual apps, coding helpers, and agentic workflows that want strong open-weight performance at a low price.

Chat

Llama 3.2 11B Vision Instruct

meta-llama/llama-3.2-11b-vision-instruct

Llama 3.2 11B Vision Instruct is Meta's multimodal model that processes both text and images with 11 billion parameters. It excels at visual recognition, image reasoning, captioning, and answering questions about images.

Chat

Llama 3.2 1B Instruct

meta-llama/llama-3.2-1b-instruct

Llama 3.2 1B Instruct is Meta's ultra-lightweight 1 billion parameter model designed for edge and mobile devices. It supports 128K context and handles summarization, instruction following, and rewriting tasks locally.

Chat

Llama 3.2 3B Instruct

meta-llama/llama-3.2-3b-instruct

Llama 3.2 3B Instruct is a compact 3 billion parameter model optimized for on-device use cases with 128K context support. It outperforms comparable models on instruction following, summarization, and tool-use tasks.

Chat

Llama 3.1 405B (base)

meta-llama/llama-3.1-405b

Llama 3.1 405B is Meta's flagship open-source large language model with 405 billion parameters, supporting 128K context length and 8 languages. It offers capabilities comparable to leading closed models for advanced reasoning, coding, and multilingual tasks.

Chat

Llama 3.1 405B Instruct

meta-llama/llama-3.1-405b-instruct

Llama 3.1 405B Instruct is the instruction-tuned version of Meta's largest open model, optimized for multilingual dialogue, tool use, and complex reasoning. It supports 8 languages with 128K context and serves as a foundation for enterprise-level AI applications.

Chat

Llama 3.1 70B Instruct

meta-llama/llama-3.1-70b-instruct

Llama 3.1 70B Instruct is a multilingual 70 billion parameter model with 128K context length, optimized for dialogue, tool use, and coding tasks. It balances strong performance with resource efficiency across 8 supported languages.

Chat

Llama 3.1 8B Instruct

meta-llama/llama-3.1-8b-instruct

Llama 3.1 8B Instruct is Meta's efficient 8 billion parameter multilingual model supporting 128K context and 8 languages. It's ideal for resource-constrained deployments requiring summarization, classification, and translation capabilities.

Chat

Llama 3.1 8B

meta-llama/llama-3.1-8b

Llama 3.1 8B is an 8-billion-parameter instruction-tuned chat model from Meta, released in July 2024 as the smallest member of the Llama 3.1 family. It handles dialogue in eight languages (English, German, French, Italian, Portuguese, Hindi, Spanish, and Thai) and was trained for tool use. Meta reports 73.0 on MMLU (0-shot, CoT), 84.5 on GSM8K, and 72.6 on HumanEval, a large improvement over Llama 3 8B's 60.4 on HumanEval, and positions it slightly ahead of Gemma 2 9B and Mistral 7B in its size class. At a few cents per million tokens, it fits high-volume work such as chat assistants, summarization, classification, and translation, where a larger model's cost is not justified.

Chat

Llama 3 70B Instruct

meta-llama/llama-3-70b-instruct

Llama 3 70B Instruct is a 70 billion parameter instruction-tuned language model from Meta, optimized for dialogue and assistant-like chat in English. It uses an optimized transformer architecture with grouped-query attention and was trained on over 15 trillion tokens.

Chat

Llama 3 8B Instruct

meta-llama/llama-3-8b-instruct

Llama 3 8B Instruct is Meta's compact 8 billion parameter instruction-tuned model for dialogue use cases in English. It offers strong performance on common benchmarks while being more efficient to deploy than its larger sibling.

Chat

LlamaGuard 2 8B

meta-llama/llama-guard-2-8b

Llama Guard 2 8B is Meta's 8 billion parameter safety classifier built on Llama 3, designed to moderate both user prompts and AI responses. It classifies content across 11 hazard categories based on the MLCommons taxonomy.

Chat

Meta Llama 3 8B Instruct Lite

meta-llama/meta-llama-3-8b-instruct-lite

Meta Llama 3 8B Instruct Lite is a cost-optimized serving tier for Meta's Llama 3 8B Instruct, an 8-billion-parameter open-weight chat model pretrained on over 15 trillion tokens. Despite its small size, it delivers strong reasoning, coding, and general-purpose conversation. On published benchmarks it scores around 67.4 on MMLU, 79.6% on GSM8K math, and 62.2% on HumanEval coding, making it highly competitive among models in its tier. The Lite tier offers the same quality at a lower per-token price, making it a great fit for developers who need fast, affordable responses for chat assistants, summarization, classification, and lightweight coding tasks at scale.

Frequently Asked Questions

What is this Meta Llama API about?

The Meta Llama API gives you access to models for AI chat. Through Puter.js, you can start using Meta Llama models instantly with zero setup or configuration.

Which Meta Llama models can I use?

Puter.js supports a variety of Meta Llama models, including Muse Glimmer 30B, Muse Spark 1.2, Muse Spark 1.1, and more. Find all AI models supported by Puter.js in the AI model list.

How much does it cost?

With the User-Pays model, users cover their own AI costs through their Puter account. This means you can build apps without worrying about infrastructure expenses.

What is Puter.js?

Puter.js is a JavaScript library that provides access to AI, storage, and other cloud services directly from a single API. It handles authentication, infrastructure, and scaling so you can focus on building your app.

Does this work with React / Vue / Vanilla JS / Node / etc.?

Yes — the Meta Llama API through Puter.js works with any JavaScript framework, Node.js, or plain HTML. Just include the library and start building. See the documentation for more details.