Skip to content
  • Models
  • Rankings
  • Ori
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Ori
  • Collections
  • Providers
  • Tools
  • Pricing
  • Business
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Trust Center
  • Support
  • Works With OR
  • Data
  • Brand

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status
  • AI Site Map

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube
Favicon for PrimeIntellect

Prime Intellect

Browse models provided by Prime Intellect

3 models

Tokens processed on OpenRouter

  • Favicon for cloudflare
    Cloudflare: Clef FlashClef Flash

    Clef-flash is the fast 9B member of Cloudflare's open-source Clef decision model family, a fine-tune of Qwen3.5-9B served on Workers AI. It turns a state (text or structured JSON) plus a schema of typed questions into decisions, returning a calibrated probability for every allowed option of every question in a single forward pass instead of generating tokens. Use it for low-latency classification, routing, scoring, and guardrails through the Decisions API. Note: Workers AI currently truncates long text state to roughly the first 2K tokens, so content beyond that is not read; images are counted separately.

    by cloudflareOct 1, 202616K context$0.021/M input tokens$0/M output tokens
  • Favicon for cloudflare
    Cloudflare: ClefClef

    Clef is Cloudflare's open-source 27B multimodal decision model, a fine-tune of Qwen3.8-27B served on Workers AI. It turns a state (text or structured JSON) plus a schema of typed questions into decisions, returning a calibrated probability for every allowed option of every question in a single forward pass instead of generating tokens. Use it for classification, routing, scoring, guardrails, and agentic control flow through the Decisions API. Note: Workers AI currently truncates long text state to roughly the first 2K tokens, so content beyond that is not read; images are counted separately.

    by cloudflareOct 1, 202616K context$0.042/M input tokens$0/M output tokens
  • Favicon for z-ai
    Z.ai: GLM 5.3GLM 5.3

    GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves on GLM-5.2 in coding and in the balance between performance and token efficiency. Reasoning is always on and cannot be disabled. Reasoning efforts low, high, and max are supported; max is the default.

    by z-aiAug 18, 20261.05M context$1.40/M input tokens$4.40/M output tokens