Model Selection
By default, NinjaCat Agents are driven by Claude Sonnet 4.6, Anthropic's best model for complex agents and coding. However, depending on your Agent’s specific needs—whether for faster response times or enhanced reasoning—you can choose an alternative model.NinjaCat will adjust the default model when newer, better models come out after spending time testing that the output is as good or better than the prior default, and if it is more cost efficient. When the default is changed, it only impacts brand new Agents. Users will need to adjust the model for existing Agents if they would like to. NinjaCat will not adjust the model for existing agents to avoid disrupting an already perfectly crafted Agent (sometimes a model change could require tweaks to the prompt).
Available Model Options in NinjaCat
Anthropic Models
| Model | Release Date | Context Window | Input $ / 1M | Output $ / 1M | Notes |
|---|---|---|---|---|---|
| Claude Sonnet 5 | Jul 2026 | ~200K | ~$3 | ~$15 | Newest Sonnet-tier model. Opt-in per agent. Pricing shown at general availability; introductory pricing applies through Aug 31, 2026. |
| Claude Sonnet 5 - Thinking | Jul 2026 | ~200K | ~$3 | ~$15 | Adaptive-thinking variant of Claude Sonnet 5 for deeper reasoning. Opt-in per agent. |
| Claude Fable 5 | Jun 2026 | ~1M | ~$10 | ~$50 | Newly restored on 7/1 following Anthropic's reinstatement announcement |
| Claude Opus 4.8 | Jun 2026 | ~1M | ~$5 | ~$25 | Latest Opus model from Anthropic; sharper reasoning and stronger performance across complex tasks. |
| Claude Opus 4.7 | Apr 2026 | ~1M | ~$5 | ~$25 | Latest Opus model; expanded ~1M context window |
| Claude Opus 4.7 - Thinking | Apr 2026 | ~1M | ~$5 | ~$25 | Enhanced reasoning variant of Opus 4.7; uses more output tokens due to adaptive thinking |
| Claude Opus 4.6 | Feb 2026 | ~200K | ~$5 | ~$25 | Superseded by Claude Opus 4.7; although costs are the same, less tokens are used on 4.6 so should still be used for less complex tasks. |
| Claude Opus 4.6 - Thinking | Feb 2026 | ~200K | ~$5 | ~$25 | Enhanced reasoning variant of Opus 4.6 |
| Claude Sonnet 4.6 (Default) | Feb 2026 | ~200K | ~$3 | ~$15 | New default for all newly created Agents. Best-balanced Claude model — strong performance at moderate cost |
| Claude Sonnet 4.6 - Thinking | Feb 2026 | ~200K | ~$3 | ~$15 | Enhanced reasoning variant of Sonnet 4.6 |
| Claude Haiku 4.5 | Oct 2025 | ~200K | ~$1 | ~$5 | Fastest & most affordable Claude option |
| Claude Haiku 4.5 - Thinking | Oct 2025 | ~200K | ~$1 | ~$5 | Enhanced reasoning variant of Haiku 4.5 |
OpenAI Models
| Model | Release Date | Context Window | Input $ / 1M | Output $ / 1M | Notes |
|---|---|---|---|---|---|
| GPT-5.6 Sol (New) | Aug 2026 | ~1M | ~$5.00 | ~$30 | Most capable GPT-5.6 tier; built for demanding coding and analysis tasks. Standard (instant) mode only — reasoning/thinking mode planned for future update. |
| GPT-5.6 Terra (New) | Aug 2026 | ~1M | ~$2.00 | ~$12 | Balanced everyday performance at a lower cost than Sol. Standard (instant) mode only — reasoning/thinking mode planned for future update. |
| GPT-5.6 Luna (New) | Aug 2026 | ~1M | ~$0.20 | ~$1.20 | Fast and lightweight; optimized for high-volume, simpler tasks. Most cost-efficient model option in NinjaCat — roughly 5× cheaper than Claude Haiku on both input and output. Standard (instant) mode only — reasoning/thinking mode planned for future update. |
| GPT-5.5 (New) | Jun 2026 | ~1M | ~$5.00 | ~$30 | OpenAI's latest model; improvements to speed and instruction |
| GPT-5.4 - Instant (New) | Mar 2026 | ~1M | ~$2.50 | ~$15 | 33% fewer errors vs GPT-5.2, massive 1M context window |
| GPT-5.4 - Thinking (New) | Mar 2026 | ~1M | ~$2.50 | ~$15 | Enhanced reasoning variant of GPT-5.4; most token-efficient reasoning model |
| GPT-5.2 - Thinking | Oct 2025 | ~400K | ~$1.75 | ~$14 | Strong reasoning + long context. Default reassignment for agents on removed OpenAI models |
| GPT-5.2 - Instant | Oct 2025 | ~400K | ~$1.75 | ~$14 | Faster, lower-latency variant of GPT-5.2 |
| GPT-5 Mini | Aug 2025 | ~400K | ~$0.25 | ~$2 | Cost-efficient; good for well-defined tasks at lower cost |
| GPT-5 Nano | Aug 2025 | ~400K | ~$0.05 | ~$0.40 | Cheapest & fastest GPT-5 variant; great for summarization and classification workloads |
GPT-5.6 models — standard mode only: GPT-5.6 Sol, Terra, and Luna run in standard (instant) mode today. Reasoning/thinking mode is not yet available for these models and is planned for a future update.
Google Models
| Model | Release Date | Context Window | Input $ / 1M | Output $ / 1M | Notes |
|---|---|---|---|---|---|
| Gemini 3.5 Flash (New) | Jun 2026 | ~1M | ~$1.50 | ~$9.00 | Google's efficiency-optimized frontier model designed for high-volume, agentic, and coding workloads |
| Gemini 3.1 Pro - Low | Feb 2026 | ~1M | ~$2.00 | ~$12.00 | Google's latest Pro model; improved reasoning, multimodal, and agentic capabilities |
| Gemini 3.1 Pro - High | Feb 2026 | ~1M | ~$2.00 | ~$12.00 | Higher reasoning effort variant of Gemini 3.1 Pro |
| Gemini 3.1 Flash Lite - Low | Feb 2026 | ~1M | ~$0.25 | ~$1.50 | Most cost-efficient Google model; optimized for high-volume agentic tasks |
| Gemini 3.1 Flash Lite - High | Feb 2026 | ~1M | ~$0.25 | ~$1.50 | Higher reasoning effort variant of Gemini 3.1 Flash Lite |
| Gemini 3 Pro - Low | Nov 2025 | ~1M | ~$2.00 | ~$12.00 | Best-in-class reasoning & multimodal from Google; massive context window |
| Gemini 3 Pro - High | Nov 2025 | ~1M | ~$2.00 | ~$12.00 | Higher reasoning effort variant |
| Gemini 3 Flash - Low | Dec 2025 | ~1M | ~$0.50 | ~$3.00 | Fast, efficient; combines Gemini 3 Pro reasoning with Flash-level latency and cost |
| Gemini 3 Flash - High | Dec 2025 | ~1M | ~$0.50 | ~$3.00 | Higher reasoning effort Flash variant |
Fireworks AI Models
| Model | Release Date | Context Window | Input $ / 1M | Output $ / 1M | Notes |
|---|---|---|---|---|---|
| GLM-5.2 (Default) | Jun 2026 | ~128K | ~$1.40 | ~$4.40 | Strong agentic and long-horizon task performance; well-suited to complex multi-step workflows and software-engineering tasks. Default Fireworks provider model. |
| DeepSeek V4 Pro | Apr 2026 | ~128K | ~$1.74 | ~$3.48 | Full-capability flagship for very long document analysis or extended reasoning; cost-efficient frontier performance. |
| DeepSeek V4 Flash | Apr 2026 | ~128K | ~$0.14 | ~$0.28 | Budget-friendly variant of V4 — significantly cheaper than Pro with fast responses; ideal for high-volume or latency-sensitive workloads. |
| Kimi K2.7 Code | Jun 2026 | ~128K | ~$0.95 | ~$4.00 | Specialized for coding and agentic coding tasks with strong token efficiency for agent workflows; supports image input (vision). |
| Qwen 3.7 Plus | Jun 2026 | ~128K | ~$0.50 | ~$3.00 | Cost-efficient agent model with a hybrid thinking mode (reasoning on/off); good for structured outputs and tool use; supports image input (vision). |
| MiniMax M3 | Jun 2026 | ~128K | ~$0.30 | ~$1.20 | Native multimodality with large-context reasoning; designed for long-context reasoning and agentic tool use; supports image input (vision). |
Note: In the Agent Builder, some models offer both a standard and a "Thinking" variant. The Thinking variant supports deeper reasoning but may come with higher cost and latency. For most agents, the standard variant is recommended unless your use case requires complex multi-step reasoning.
How to Choose the Right Model
AI models are continuously improving — what is "best" today may be surpassed in weeks or months. NinjaCat will continue evaluating and adding models that demonstrate better intelligence, efficiency, or performance.
General guidance:
- For most agents: Claude Sonnet 4.6 (default) is the best starting point — strong performance at reasonable cost.
- For the most complex, high-effort, or long-horizon tasks: Claude Opus 4.8 or Opus 4.7 — Anthropic's highest-capability models available on the platform. (Note: Claude Fable 5 is currently unavailable due to a government directive — see the Anthropic models table above.)
- For complex reasoning or coding tasks: Claude Opus 4.7, GPT-5.4 - Thinking, GPT-5.2 - Thinking, or GPT-5.6 Sol (standard mode).
- For speed or cost-sensitive tasks: Claude Haiku 4.5, GPT-5.6 Luna (most cost-efficient option on the platform — ~5× cheaper than Claude Haiku), GPT-5 Nano, Gemini 3.1 Flash Lite, or Gemini 3 Flash.
- For balanced everyday performance at a lower cost: GPT-5.6 Terra is a good middle-ground option within the GPT-5.6 family.
- For large context windows: OpenAI GPT-5.4 series (~1M), GPT-5.6 series (~1M), or Google Gemini series (~1M).
For the latest information from each provider, see their documentation:
Anthropic Claude Models OpenAI GPT-5 Prompting Guide Google Gemini
Note: When switching between models, prompt adjustments may be required to maintain optimal Agent performance. We will provide further guidance on prompt modifications as we continue testing and learning.
Updated 7 days ago