Frontier-scale open models.

Copy this into your agent to get started

One inference stack

Use the hardware the workload needs.

01 · Hosted

Wally

Frontier-scale open models behind an OpenAI-compatible API. Wally handles browser sign-in and keeps hosted execution explicit. A request only leaves your machine when you say so.

Explore Wally

02 · Accelerated locally

MetalRT & QHexRT

NPU engines for the silicon the device already ships. Every published result names its hardware and method.

Explore accelerators

03 · Build anywhere

Open-source SDKs

Six bindings on one C++ core, so behaviour does not drift between platforms.

Read the developer docs

Console

Sign in, buy credits, run.

Usage and spend, live, in one place.

Open the console

Start with Wally.