How to Put a Small AI in a Device: Small LLMs on Phones, Embedded Boards and Tiny Hardware, and the Parts They Need
8 October 2026
I usually run local LLMs on a mini PC. In September, Meta announced an AI agent called Muse, along with a keychain-sized device for talking to it. According to Meta’s official announcement, Muse runs on a dedicated computer in the cloud. So what hardware would you need to put a small AI into a device you build yourself, without relying on the cloud?
This article summarizes what I found, mostly from official documentation, in this order: smartphones, embedded boards, and devices smaller than a phone. It also covers the setup where a small device connects to a local LLM at home. Mini PCs are out of scope.
This reflects information as of 3 October 2026. I have not run any of the devices covered. This is a summary of published material, not a record of my own measurements. Next to each figure I note whether it comes from an official source, press reports, or a third party.
- 1. What I looked into and how I checked
- 2. Where does Meta’s Muse run?
- 3. How big is the AI that runs inside a phone?
- 4. Can a small LLM run on an embedded board?
- 5. What runs on devices smaller than a phone?
- 6. Connecting a small device to a local LLM at home
- 7. What parts do you need if you build one yourself?
- 8. What are the applications?
- 9. What I could not confirm this time
- 10. Where would I start if putting AI in a device?
What I looked into and how I checked
I looked at where the AI lives when you put a small AI (a small LLM) into a device, split into three places.
| Where the AI lives | What that means | Covered here |
|---|---|---|
| 1. On the device | The model sits in the device’s own memory and runs on the device alone | Smartphones, embedded boards, devices smaller than a phone |
| 2. A server at home | The device only sends audio and the like; the model runs on a PC or server at home | Small device + local LLM at home (touched on as “where the brain lives") |
| 3. The cloud | The device only communicates; the model runs in the provider’s cloud | Checked using Meta’s Muse as the example |
What I checked is a rough guide to the size of model that can run, and the parts needed on the device for it (memory, an AI chip, power).
Where does Meta’s Muse run?
Meta announced Muse on 8 September 2026. Meta’s official announcement (8 September 2026) says Muse runs on a dedicated cloud computer and is isolated so that other people’s agents cannot reach it. I could not find any statement that the agent runs on the device side (the keychain device, a phone or glasses).
About the keychain device “Muse Charm" shown at Meta Connect on 23 September, what Meta says officially is that it is a device for talking to Muse and that it is planned to ship at the end of the year. The roughly 2-inch screen, front and rear cameras and built-in 5G come from press reports (secondary reports derived from Bloomberg) and are not confirmed by Meta. Size, battery, chip and price have not been disclosed.
For glasses, Meta explains “Private Processing" in an official engineering blog post dated 23 September 2026. Many voice commands, such as calls and replying to messages, are completed on the device. Tasks that need a larger model, such as context search and long-term memory, are said to run in encrypted virtual machines in the cloud. The post does not mention Muse by name, so I could not confirm whether Private Processing is used for Muse.
From this, a device like Muse Charm looks like an input/output terminal for a cloud AI. The built-in 5G, which lets it connect without a phone, also fits a design that does not put an LLM on the device (this is my inference; Meta has not said whether any inference runs on the device).
How big is the AI that runs inside a phone?
Of the three places, on-device AI on smartphones has the best official documentation.
| Implementation | Model size | Official description | Source |
|---|---|---|---|
| Apple (on-device model, iOS 26 generation) | About 3B (3 billion parameters) | Weights quantized to 2 bits. Supports iPhone 15 Pro and later. RAM size is not stated on the official page | Official (June and July 2025) |
| Apple (AFM 3 Core, iOS 27 generation) | 3B | Supports iPhone 16 and later, plus 15 Pro and 15 Pro Max. No official speed figure | Official (8 June 2026) |
| Google Gemma 3n (open model) | Effective about 2B (E2B) / about 4B (E4B) | Memory guideline: 2GB for E2B, 3GB for E4B. Raw parameters are over 5B and 8B | Official (June 2025) |
| Google Gemini Nano (AICore) | Not stated in official documentation | Reported that Gemini Intelligence features require, for example, 12GB or more of RAM (quoting a footnote on a Google official page) | Press (May 2026) |
I could not find an official document that summarizes “how many B of model runs with how many GB of RAM". Third-party aggregator blogs have rough estimates, but they are not measurements, and I only saw search-result summaries without reading the full text, so I have not included them.
As reasons to run on the device, the official Android documentation lists that prompts are not sent to a server so sensitive data stays on the device, and that it works without an internet connection. The person in charge of the summarization feature in Google’s Recorder app gives privacy, low latency and no need for a network as reasons (August 2024).
Can a small LLM run on an embedded board?
On boards built into devices, memory size and memory speed (bandwidth) seem to be the big constraints, especially in the generation phase. Here are five types I could confirm from official and third-party material.
| Board | Memory | AI performance | Power | Official LLM support / measurements | Price (date checked) |
|---|---|---|---|---|---|
| Raspberry Pi 5 + AI HAT+ 2 (Hailo-10H) | 8GB on the AI HAT side | 40 TOPS (INT4) | Hailo-10H typically 2.5W (official) | At launch, five supported models in the 1B-1.5B class (official) | $200 on the official page, $130 in the official launch news (they disagree) |
| NVIDIA Jetson Orin Nano Super | 8GB, 102GB/s | 67 TOPS (sparse INT8) | 7-25W | Measured in the official blog: Llama 3.1 8B 19.14 tok/s, Qwen2.5 7B 21.75 tok/s, Llama 3.2 3B 43.07 tok/s (MLC, INT4) | $249 in the US, 106,920 yen in Japan including tax (1 October 2026) |
| Rockchip RK3588 family (Radxa ROCK 5B+ etc.) | Up to 32GB | NPU 6 TOPS (official) | Actual consumption not found | Official RKLLM supports Qwen, Llama, Gemma and others | Not found |
| Rockchip RK1820 / RK1828 (LLM coprocessors) | Dedicated DRAM 2.5GB / 5GB (CNX Software report) | 20 TOPS (INT8) (same report) | Not found | Targets 3B / 7B models, reported at 59-180 tok/s (CNX Software). I could not read the official datasheet | Dev kits $889-$1,029 (reported) |
| Arduino UNO Q (Qualcomm QRB2210) | 2GB or 4GB | No TOPS figure officially | Not found | The official page only says it can host “small LLMs". No model names or speeds | $59 at launch for the 4GB version. A US store listing seen in search results on 1 October 2026 showed $79 (shown as out of stock) |
In the table, TOPS (a rough measure of an AI chip’s compute) cannot be compared side by side. Hailo-10H’s 40 TOPS is INT4 and Jetson’s 67 TOPS is sparse INT8, figures under different conditions.
Is TOPS not the whole story?
In CNX Software’s measurements, the Raspberry Pi 5’s CPU generated faster than the 40 TOPS AI HAT+ 2. For example, on Llama 3.2 3B the CPU got 4.78 tok/s and the AI HAT+ 2 got 2.60 tok/s (third-party measurement). CNX Software quotes Raspberry Pi’s Eben Upton as saying the Pi 5 and the Hailo-10 have the same memory system, LPDDR4X-4267, and that memory bandwidth is the limiting factor.
NVIDIA’s official blog also explains that in the generation phase of an LLM, the speed of reading weights and the like from memory sets the latency, not the speed of computation. However, I could not find a statement from the Raspberry Pi, Rockchip or NVIDIA official sources saying in general terms that “memory bandwidth is the limiting factor".
What runs on devices smaller than a phone?
Glasses, watches and earbuds
| Device | Model said to run on the device | Source |
|---|---|---|
| Smart glasses (Qualcomm Snapdragon AR1+ Gen 1) | A small LLM of up to 1B (1 billion) parameters. A live demo ran Llama 1B on the device (June 2025) | Press (quoting Qualcomm). No shipping date for a mass-market product found |
| Smartwatch (Qualcomm Snapdragon Wear Elite) | Up to 2B (2 billion) parameters (March 2026) | Press. I could not retrieve the body of Qualcomm’s official release |
| Earbuds and hearing aids | No example found where an official source says an LLM runs on the device | – |
For Meta’s Ray-Ban Meta glasses, I could not retrieve Meta’s official spec sheet, so I could not confirm whether there is an LLM on the device. For Ray-Ban Display, Meta’s official help says that using Meta AI requires a smartphone with the Meta AI app and Wi-Fi.
Microcontrollers
On microcontrollers (small control chips), the examples I found were mostly tiny models built for a narrow purpose, such as TinyStories.
| Chip | Model run | Speed | Source |
|---|---|---|---|
| ESP32-S3 | 260K-parameter TinyStories (a model that only writes short children’s stories) | 19.13 tok/s | Third party (GitHub, author’s own claim) |
| ESP32-S3 (about $8 dev board) | 28.9M-parameter TinyStories (4-bit) | About 9.5 tok/s | Third party (GitHub, author’s own claim). It cannot answer questions or follow instructions |
| Raspberry Pi Pico 2 + microSD | Qwen3-0.6B (Q4_0, 321MB) | About 19.4 seconds per token | Third party (GitHub). Weights are read from the SD card |
| Arm Ethos-U85 (FPGA prototype) | 15M-parameter Tiny Llama2 | 7.5-8 tok/s | Official (Arm, November 2024) |
Within what I checked, I found no example of a general-purpose small LLM of several hundred million parameters or more running on a microcontroller alone at a practical speed.
Very small models
As candidates for models to place on a device, here are small models I could confirm from official model cards and distribution sites. File sizes are mostly actual sizes I checked in third-party distribution repositories, not official values from the developers.
I confirmed the context length of Gemma 3 270M (32K) on the official Hugging Face model card (3 October 2026).
| Model | Parameters | Context length | Size after 4-bit quantization | Source |
|---|---|---|---|---|
| Gemma 3 270M | 270M | 32K (32,768 tokens) | 253.1MB (Q4_K_M) | Official (specs) + third party (size) |
| SmolLM2-135M | About 135M | 8K (8,192 tokens) | 105.5MB (Q4_K_M) | Official + third party |
| SmolLM2-360M | 360M | 8K (8,192 tokens) | 270.6MB (Q4_K_M) | Official + third party |
| LFM2-350M | About 350M | 32K (32,768 tokens) | 229.3MB (Q4_K_M) | Official (Liquid’s distribution) |
| Qwen3-0.6B | 600M | 32K (32,768 tokens) | 639.4MB (Q8_0, 8-bit) | Official |
| Llama 3.2 1B | 1.23B | 128K (128,000 tokens) | 807.7MB (Q4_K_M) | Official + third party |
Some models are pitched for narrowly scoped use. For example, the official description says Gemma 3 270M is meant to be fine-tuned for things like sentiment analysis, entity extraction and classification. Google writes that with the 4-bit version of Gemma 3 270M on a Pixel 9 Pro, 25 conversations used 0.75% of the battery. As for Japanese support, the only one I could confirm was in the official description of LFM2-350M.
Connecting a small device to a local LLM at home
Running a “small device + AI" setup like Muse on a home PC instead of the cloud exists as an official mechanism: Assist, the voice assistant of Home Assistant (open-source home automation software).
| Role | Where it lives | Official description |
|---|---|---|
| Microphone and speaker | A small device (ESP32-based Voice PE, ATOM Echo, ESP32-S3-BOX-3, etc.) | The device streams microphone audio to Home Assistant, and processing happens on the Home Assistant side (official ESPHome) |
| Speech recognition (Whisper) | Server at home | About 8 seconds on a Raspberry Pi 4, under 1 second on an Intel NUC (official Home Assistant) |
| LLM | Server at home (Ollama, etc.) | Ollama can be connected as an external server over the network (official) |
| Speech synthesis (Piper) | Server at home | With streaming support, the time until speech starts went from “over 5 seconds" to “about 0.5 seconds" (stated for both Piper and cloud; official Home Assistant blog, 22 October 2025; third-party figures with decimals are 0.56 / 0.51 seconds). Audio generation starts as soon as the LLM’s first sentence appears |
Home Assistant’s official documentation says local models that use 8GB or more of VRAM have come close to cloud models (September 2025). On the other hand, small models make mistakes more easily, and it is recommended to expose fewer than 25 devices for the LLM to control. LLM integration through the Assist API is described as experimental.
As a wearable example, OMI (an open-source AI pendant) officially publishes steps for a self-hosted backend. But it assumes cloud APIs such as Firebase and OpenAI, and I found no steps for swapping in a local LLM. The only official Ollama support I found was for image descriptions on OMI’s ESP32-S3 glasses (omiGlass).
Watch out for older implementations too. Open Interpreter’s 01 has a local mode that uses Ollama, but its last update was 1 November 2024, about 11 months ago. Home Assistant’s old satellite implementation wyoming-satellite was archived in January 2026, and its successor is said to be Linux Voice Assistant.
For reaching a home server from outside, Tailscale officially explains that it connects devices directly where possible, that a direct connection has the lowest latency, and that relays are slower. However, I found no official figures for the effect on battery or data usage, and no finished official product for “small wearable to a local LLM at home". Connecting from outside the home is a combination that is conceivable, not an established setup.
What parts do you need if you build one yourself?
Will the model fit? Memory capacity
A model’s weight size is the parameter count times the bytes per parameter. NVIDIA’s official blog gives the example that holding 7B at 16 bits (FP16) takes about 14GB. llama.cpp and others explain that quantization to 1.5-8 bits reduces memory. Google says Gemma 3n runs with accelerator memory on the scale of 2GB and 3GB even though the raw parameters are over 5B and 8B (this is about working memory, not download size).
Will it be fast enough? Memory bandwidth
As NVIDIA’s explanation in the earlier section says, the generation phase is dominated by how fast memory can be read. Some papers say that AI chips (NPUs) are strong at reading in the input but that a CPU can be better for generation (arXiv; I have not read the original). Papers also seem to report that phones slow down from heat when used continuously, but I have not read the originals, so I give no figures.
Examples of parts on the device side
| Setup | Example device-side parts | Source |
|---|---|---|
| Input/output device that connects to the cloud or home | ESP32-S3 (Xtensa LX7, 2 cores, 240MHz, Wi-Fi and Bluetooth LE), microphone, speaker, battery. The Seeed XIAO ESP32S3 Sense has a built-in camera and microphone, 8MB PSRAM, 8MB flash, 21 x 17.8mm | Official |
| Putting a small LLM on the device | Raspberry Pi Zero 2 W (Cortex-A53, 4 cores, 1GHz, 512MB, 65 x 30mm). With SmolLM2-360M, about 2.99 tok/s generation and about 40 seconds per response (author’s own claim, not replicated) | Official (specs) + third party (measurement) |
The one DIY example I confirmed that runs entirely on the device (Raspberry Pi Zero 2 W) took about 40 seconds per response. I found no example with a camera that answers instantly the way Meta’s does.
What are the applications?
I limited the examples of small LLMs built into devices to those with a source, and separated shipped products from prototypes.
| Field | Example | Model info / performance | Status | Source |
|---|---|---|---|---|
| Phone / recording | Pixel Recorder summaries (Gemini Nano) | Not stated officially | Shipped (August 2024) | Official |
| Phone / OS | Apple Intelligence on-device model | About 3B | Shipped | Official |
| PC | Phi Silica on Windows Copilot+ PCs | Context length 4,000 tokens, 230 milliseconds to the first token | Shipped | Official (December 2024) |
| Translation | Galaxy AI Interpreter (interprets without a data connection once language packs are installed) | Model used not found | Shipped | Official |
| In-car | Cerence CaLLM Edge (vehicle controls and place search without a connection) | 3.8B | Announced. No vehicle that ships with it found (November 2024) | Press (reprint of a press release) |
| Industrial equipment | Rockwell FactoryTalk Design Studio Copilot (described as able to run on control panels and the like) | 9B class | Announced (November 2025). Whether it has shipped not found | Official press release (13 November 2025. Described as built on Nemotron-Nano-9B-v2, fine-tuned on FactoryTalk data and run on operator panels and dedicated devices. I confirmed the body through press quotes) |
| Glasses | Glasses with Snapdragon AR1+ Gen 1 | 1B | Prototype / demo (June 2025) | Press |
| Parts / robots | M5Stack Module LLM (fully offline) | Qwen2.5-0.5B. NPU 3.2 TOPS, 4GB RAM, about 1.5W | Product (on sale) | Official |
| Care / mental health support | Research on a hospital robot (run locally; reason given is GDPR compliance) / research on an offline mental-health support app | 13B / 1B | Research | Papers |
For mass-produced home appliances, explaining anomalies in industrial equipment, and mass-market toys, I could not confirm from primary sources any shipped product where a small LLM runs on the device.
What I could not confirm this time
- Muse Charm’s size, battery, chip and price (not disclosed by Meta). Screen, cameras and 5G are press reports only
- Whether Meta’s Ray-Ban Meta glasses have an LLM on the device (I could not retrieve Meta’s official spec sheet)
- An official list of “how many GB of RAM runs up to how many B" for phones, and Gemini Nano’s parameter count
- An example where an official source explains an LLM running on earbuds or hearing aids
- An example of a general-purpose small LLM of several hundred million parameters or more running on a microcontroller alone at a practical speed
- A finished official product connecting a wearable to a local LLM at home, and official figures for Tailscale’s battery and data usage
Where would I start if putting AI in a device?
In the setup where the device only handles input and output and the brain lives at home, you need a home PC or server separate from the device. From what I found, what the reader needs to decide is where to put the AI for their own device.
| What you want to build | Rough guide for where the AI lives | Basis |
|---|---|---|
| A smartphone app | The OS’s built-in on-device model (Apple, Gemini Nano), or a small model like Gemma 3n | Official sources state the uses (summarizing, rewriting, short conversations) and supported devices |
| A palm-sized dedicated device (with mains power) | A board in the class of Jetson Orin Nano Super: 8GB memory, 102GB/s bandwidth. The officially supported models for Raspberry Pi 5 + AI HAT+ 2 at launch include 1B-1.5B class | Official blog measurements (Jetson), third-party measurements (Raspberry Pi) |
| A small device operated by voice inside the home | The device does input and output only; the brain is a home server (Home Assistant + Ollama) | An official setup. Official figures exist only for Whisper and Piper; LLM latency is a third-party reference value |
| A battery-powered tiny device | I found no realistic example of running a general-purpose small LLM on a microcontroller alone. Connect to the cloud or home over a network | The microcontroller examples are toy-sized |
| A device used outdoors like Muse | Muse connects to the cloud. Connecting to a home server needs a combination such as a VPN | Meta official (cloud). Tailscale’s official explanation |
If you are unsure, the entry point with the most documentation is to install Ollama on a home PC and follow Home Assistant’s official steps (the Ollama integration) to talk to it from a small device. For ESP32-based devices, there are several products with official support.
Sources consulted
- Meta: official Muse announcement (8 September 2026) / Meta official engineering blog “Private Processing" (23 September 2026)
- Apple Machine Learning Research: Apple Foundation Models 2025 / the Apple Intelligence announcement of 8 June 2026
- Google Developers Blog: Gemma 3n, Gemma 3 270M / Android Developers: Gemini Nano
- Raspberry Pi: AI HAT+ 2 / CNX Software: AI HAT+ 2 review (20 January 2026), RK1820 and RK1828 (30 December 2025)
- NVIDIA Developer Blog: speeding up Jetson Orin Nano / optimizing LLM inference
- Home Assistant: Assist, the Ollama integration, official blog Voice Chapter 11 (22 October 2025)
- Hugging Face: model cards and distribution repositories / GitHub: LLM examples on ESP32 and Pico / Arm: Small Language Model on Edge (November 2024)









Discussion
New Comments
No comments yet. Be the first one!