SimpleLM

Memory is not an option for humans. Privacy isn’t either.

Same as us. On your device. Context that stays with the chat.

Get the highlights.

On device

The model wakes where you are.

No cloud round-trip to start. Intelligence loads on the seat, then stays.

Airplane mode

Flip the switch. Keep going.

Network can leave. Your assistant stays. Private by physics, not by policy.

Generation

Answers that never leave the seat.

Tokens stream on-device. Same chat, same brain, even when the tower is gone.

Swipe to explore

Memory

A million tokens.
Still on your device.

One chat can grow for months, and the context stays where you are. No cloud lease for the memory that makes it yours.

One million tokens per chat on your device with Pro

In plain terms

A local AI assistant that runs on your iPhone.

SimpleLM runs a local LLM on the device itself. There is no account to create, no server to sign in to, and no cloud round-trip for anything you type. Turn on airplane mode and it keeps working, because the model is already on the phone.

What separates it from other private AI apps is memory. Small models on a phone usually lose the thread after a few thousand tokens and start contradicting themselves. SimpleLM holds up to one million tokens in a single conversation, and we publish the measurements rather than asserting them.

1M
tokens of contextper conversation, held on the device
None
network callsno connection needed to generate
Zero
accountsnothing to sign up for, ever
2B Q4
four-bit modelrunning on Apple silicon

Where the context figure comes from

A million tokens is the claim worth checking, so we published how it was measured. The benchmarks carry RULER, NIAH and MRCR results beside published leaderboards, the engine page explains how it works, and the technical report carries the full method.

Your private assistant. On your device.