How to Put a Small AI in a Device: Small LLMs on Phones, Embedded Boards and Tiny Hardware, and the Parts They Need

本ページは広告(アフィリエイトプログラム)を含みます。詳しくはプライバシーポリシーをご覧ください。

8 October 2026

I usually run local LLMs on a mini PC. In September, Meta announced an AI agent called Muse, along with a keychain-sized device for talking to it. According to Meta’s official announcement, Muse runs on a dedicated computer in the cloud. So what hardware would you need to put a small AI into a device you build yourself, without relying on the cloud?

This article summarizes what I found, mostly from official documentation, in this order: smartphones, embedded boards, and devices smaller than a phone. It also covers the setup where a small device connects to a local LLM at home. Mini PCs are out of scope.

This reflects information as of 3 October 2026. I have not run any of the devices covered. This is a summary of published material, not a record of my own measurements. Next to each figure I note whether it comes from an official source, press reports, or a third party.

What I looked into and how I checked

I looked at where the AI lives when you put a small AI (a small LLM) into a device, split into three places.

Where the AI livesWhat that meansCovered here
1. On the deviceThe model sits in the device’s own memory and runs on the device aloneSmartphones, embedded boards, devices smaller than a phone
2. A server at homeThe device only sends audio and the like; the model runs on a PC or server at homeSmall device + local LLM at home (touched on as “where the brain lives")
3. The cloudThe device only communicates; the model runs in the provider’s cloudChecked using Meta’s Muse as the example

What I checked is a rough guide to the size of model that can run, and the parts needed on the device for it (memory, an AI chip, power).

Where does Meta’s Muse run?

Meta announced Muse on 8 September 2026. Meta’s official announcement (8 September 2026) says Muse runs on a dedicated cloud computer and is isolated so that other people’s agents cannot reach it. I could not find any statement that the agent runs on the device side (the keychain device, a phone or glasses).

About the keychain device “Muse Charm" shown at Meta Connect on 23 September, what Meta says officially is that it is a device for talking to Muse and that it is planned to ship at the end of the year. The roughly 2-inch screen, front and rear cameras and built-in 5G come from press reports (secondary reports derived from Bloomberg) and are not confirmed by Meta. Size, battery, chip and price have not been disclosed.

For glasses, Meta explains “Private Processing" in an official engineering blog post dated 23 September 2026. Many voice commands, such as calls and replying to messages, are completed on the device. Tasks that need a larger model, such as context search and long-term memory, are said to run in encrypted virtual machines in the cloud. The post does not mention Muse by name, so I could not confirm whether Private Processing is used for Muse.

From this, a device like Muse Charm looks like an input/output terminal for a cloud AI. The built-in 5G, which lets it connect without a phone, also fits a design that does not put an LLM on the device (this is my inference; Meta has not said whether any inference runs on the device).

How big is the AI that runs inside a phone?

Of the three places, on-device AI on smartphones has the best official documentation.

ImplementationModel sizeOfficial descriptionSource
Apple (on-device model, iOS 26 generation)About 3B (3 billion parameters)Weights quantized to 2 bits. Supports iPhone 15 Pro and later. RAM size is not stated on the official pageOfficial (June and July 2025)
Apple (AFM 3 Core, iOS 27 generation)3BSupports iPhone 16 and later, plus 15 Pro and 15 Pro Max. No official speed figureOfficial (8 June 2026)
Google Gemma 3n (open model)Effective about 2B (E2B) / about 4B (E4B)Memory guideline: 2GB for E2B, 3GB for E4B. Raw parameters are over 5B and 8BOfficial (June 2025)
Google Gemini Nano (AICore)Not stated in official documentationReported that Gemini Intelligence features require, for example, 12GB or more of RAM (quoting a footnote on a Google official page)Press (May 2026)

I could not find an official document that summarizes “how many B of model runs with how many GB of RAM". Third-party aggregator blogs have rough estimates, but they are not measurements, and I only saw search-result summaries without reading the full text, so I have not included them.

As reasons to run on the device, the official Android documentation lists that prompts are not sent to a server so sensitive data stays on the device, and that it works without an internet connection. The person in charge of the summarization feature in Google’s Recorder app gives privacy, low latency and no need for a network as reasons (August 2024).

Can a small LLM run on an embedded board?

On boards built into devices, memory size and memory speed (bandwidth) seem to be the big constraints, especially in the generation phase. Here are five types I could confirm from official and third-party material.

BoardMemoryAI performancePowerOfficial LLM support / measurementsPrice (date checked)
Raspberry Pi 5 + AI HAT+ 2 (Hailo-10H)8GB on the AI HAT side40 TOPS (INT4)Hailo-10H typically 2.5W (official)At launch, five supported models in the 1B-1.5B class (official)$200 on the official page, $130 in the official launch news (they disagree)
NVIDIA Jetson Orin Nano Super8GB, 102GB/s67 TOPS (sparse INT8)7-25WMeasured in the official blog: Llama 3.1 8B 19.14 tok/s, Qwen2.5 7B 21.75 tok/s, Llama 3.2 3B 43.07 tok/s (MLC, INT4)$249 in the US, 106,920 yen in Japan including tax (1 October 2026)
Rockchip RK3588 family (Radxa ROCK 5B+ etc.)Up to 32GBNPU 6 TOPS (official)Actual consumption not foundOfficial RKLLM supports Qwen, Llama, Gemma and othersNot found
Rockchip RK1820 / RK1828 (LLM coprocessors)Dedicated DRAM 2.5GB / 5GB (CNX Software report)20 TOPS (INT8) (same report)Not foundTargets 3B / 7B models, reported at 59-180 tok/s (CNX Software). I could not read the official datasheetDev kits $889-$1,029 (reported)
Arduino UNO Q (Qualcomm QRB2210)2GB or 4GBNo TOPS figure officiallyNot foundThe official page only says it can host “small LLMs". No model names or speeds$59 at launch for the 4GB version. A US store listing seen in search results on 1 October 2026 showed $79 (shown as out of stock)

In the table, TOPS (a rough measure of an AI chip’s compute) cannot be compared side by side. Hailo-10H’s 40 TOPS is INT4 and Jetson’s 67 TOPS is sparse INT8, figures under different conditions.

Is TOPS not the whole story?

In CNX Software’s measurements, the Raspberry Pi 5’s CPU generated faster than the 40 TOPS AI HAT+ 2. For example, on Llama 3.2 3B the CPU got 4.78 tok/s and the AI HAT+ 2 got 2.60 tok/s (third-party measurement). CNX Software quotes Raspberry Pi’s Eben Upton as saying the Pi 5 and the Hailo-10 have the same memory system, LPDDR4X-4267, and that memory bandwidth is the limiting factor.

NVIDIA’s official blog also explains that in the generation phase of an LLM, the speed of reading weights and the like from memory sets the latency, not the speed of computation. However, I could not find a statement from the Raspberry Pi, Rockchip or NVIDIA official sources saying in general terms that “memory bandwidth is the limiting factor".

What runs on devices smaller than a phone?

Glasses, watches and earbuds

DeviceModel said to run on the deviceSource
Smart glasses (Qualcomm Snapdragon AR1+ Gen 1)A small LLM of up to 1B (1 billion) parameters. A live demo ran Llama 1B on the device (June 2025)Press (quoting Qualcomm). No shipping date for a mass-market product found
Smartwatch (Qualcomm Snapdragon Wear Elite)Up to 2B (2 billion) parameters (March 2026)Press. I could not retrieve the body of Qualcomm’s official release
Earbuds and hearing aidsNo example found where an official source says an LLM runs on the device–

For Meta’s Ray-Ban Meta glasses, I could not retrieve Meta’s official spec sheet, so I could not confirm whether there is an LLM on the device. For Ray-Ban Display, Meta’s official help says that using Meta AI requires a smartphone with the Meta AI app and Wi-Fi.

Microcontrollers

On microcontrollers (small control chips), the examples I found were mostly tiny models built for a narrow purpose, such as TinyStories.

ChipModel runSpeedSource
ESP32-S3260K-parameter TinyStories (a model that only writes short children’s stories)19.13 tok/sThird party (GitHub, author’s own claim)
ESP32-S3 (about $8 dev board)28.9M-parameter TinyStories (4-bit)About 9.5 tok/sThird party (GitHub, author’s own claim). It cannot answer questions or follow instructions
Raspberry Pi Pico 2 + microSDQwen3-0.6B (Q4_0, 321MB)About 19.4 seconds per tokenThird party (GitHub). Weights are read from the SD card
Arm Ethos-U85 (FPGA prototype)15M-parameter Tiny Llama27.5-8 tok/sOfficial (Arm, November 2024)

Within what I checked, I found no example of a general-purpose small LLM of several hundred million parameters or more running on a microcontroller alone at a practical speed.

Very small models

As candidates for models to place on a device, here are small models I could confirm from official model cards and distribution sites. File sizes are mostly actual sizes I checked in third-party distribution repositories, not official values from the developers.

I confirmed the context length of Gemma 3 270M (32K) on the official Hugging Face model card (3 October 2026).

ModelParametersContext lengthSize after 4-bit quantizationSource
Gemma 3 270M270M32K (32,768 tokens)253.1MB (Q4_K_M)Official (specs) + third party (size)
SmolLM2-135MAbout 135M8K (8,192 tokens)105.5MB (Q4_K_M)Official + third party
SmolLM2-360M360M8K (8,192 tokens)270.6MB (Q4_K_M)Official + third party
LFM2-350MAbout 350M32K (32,768 tokens)229.3MB (Q4_K_M)Official (Liquid’s distribution)
Qwen3-0.6B600M32K (32,768 tokens)639.4MB (Q8_0, 8-bit)Official
Llama 3.2 1B1.23B128K (128,000 tokens)807.7MB (Q4_K_M)Official + third party

Some models are pitched for narrowly scoped use. For example, the official description says Gemma 3 270M is meant to be fine-tuned for things like sentiment analysis, entity extraction and classification. Google writes that with the 4-bit version of Gemma 3 270M on a Pixel 9 Pro, 25 conversations used 0.75% of the battery. As for Japanese support, the only one I could confirm was in the official description of LFM2-350M.

Connecting a small device to a local LLM at home

Running a “small device + AI" setup like Muse on a home PC instead of the cloud exists as an official mechanism: Assist, the voice assistant of Home Assistant (open-source home automation software).

RoleWhere it livesOfficial description
Microphone and speakerA small device (ESP32-based Voice PE, ATOM Echo, ESP32-S3-BOX-3, etc.)The device streams microphone audio to Home Assistant, and processing happens on the Home Assistant side (official ESPHome)
Speech recognition (Whisper)Server at homeAbout 8 seconds on a Raspberry Pi 4, under 1 second on an Intel NUC (official Home Assistant)
LLMServer at home (Ollama, etc.)Ollama can be connected as an external server over the network (official)
Speech synthesis (Piper)Server at homeWith streaming support, the time until speech starts went from “over 5 seconds" to “about 0.5 seconds" (stated for both Piper and cloud; official Home Assistant blog, 22 October 2025; third-party figures with decimals are 0.56 / 0.51 seconds). Audio generation starts as soon as the LLM’s first sentence appears

Home Assistant’s official documentation says local models that use 8GB or more of VRAM have come close to cloud models (September 2025). On the other hand, small models make mistakes more easily, and it is recommended to expose fewer than 25 devices for the LLM to control. LLM integration through the Assist API is described as experimental.

As a wearable example, OMI (an open-source AI pendant) officially publishes steps for a self-hosted backend. But it assumes cloud APIs such as Firebase and OpenAI, and I found no steps for swapping in a local LLM. The only official Ollama support I found was for image descriptions on OMI’s ESP32-S3 glasses (omiGlass).

Watch out for older implementations too. Open Interpreter’s 01 has a local mode that uses Ollama, but its last update was 1 November 2024, about 11 months ago. Home Assistant’s old satellite implementation wyoming-satellite was archived in January 2026, and its successor is said to be Linux Voice Assistant.

For reaching a home server from outside, Tailscale officially explains that it connects devices directly where possible, that a direct connection has the lowest latency, and that relays are slower. However, I found no official figures for the effect on battery or data usage, and no finished official product for “small wearable to a local LLM at home". Connecting from outside the home is a combination that is conceivable, not an established setup.

What parts do you need if you build one yourself?

Will the model fit? Memory capacity

A model’s weight size is the parameter count times the bytes per parameter. NVIDIA’s official blog gives the example that holding 7B at 16 bits (FP16) takes about 14GB. llama.cpp and others explain that quantization to 1.5-8 bits reduces memory. Google says Gemma 3n runs with accelerator memory on the scale of 2GB and 3GB even though the raw parameters are over 5B and 8B (this is about working memory, not download size).

Will it be fast enough? Memory bandwidth

As NVIDIA’s explanation in the earlier section says, the generation phase is dominated by how fast memory can be read. Some papers say that AI chips (NPUs) are strong at reading in the input but that a CPU can be better for generation (arXiv; I have not read the original). Papers also seem to report that phones slow down from heat when used continuously, but I have not read the originals, so I give no figures.

Examples of parts on the device side

SetupExample device-side partsSource
Input/output device that connects to the cloud or homeESP32-S3 (Xtensa LX7, 2 cores, 240MHz, Wi-Fi and Bluetooth LE), microphone, speaker, battery. The Seeed XIAO ESP32S3 Sense has a built-in camera and microphone, 8MB PSRAM, 8MB flash, 21 x 17.8mmOfficial
Putting a small LLM on the deviceRaspberry Pi Zero 2 W (Cortex-A53, 4 cores, 1GHz, 512MB, 65 x 30mm). With SmolLM2-360M, about 2.99 tok/s generation and about 40 seconds per response (author’s own claim, not replicated)Official (specs) + third party (measurement)

The one DIY example I confirmed that runs entirely on the device (Raspberry Pi Zero 2 W) took about 40 seconds per response. I found no example with a camera that answers instantly the way Meta’s does.

What are the applications?

I limited the examples of small LLMs built into devices to those with a source, and separated shipped products from prototypes.

FieldExampleModel info / performanceStatusSource
Phone / recordingPixel Recorder summaries (Gemini Nano)Not stated officiallyShipped (August 2024)Official
Phone / OSApple Intelligence on-device modelAbout 3BShippedOfficial
PCPhi Silica on Windows Copilot+ PCsContext length 4,000 tokens, 230 milliseconds to the first tokenShippedOfficial (December 2024)
TranslationGalaxy AI Interpreter (interprets without a data connection once language packs are installed)Model used not foundShippedOfficial
In-carCerence CaLLM Edge (vehicle controls and place search without a connection)3.8BAnnounced. No vehicle that ships with it found (November 2024)Press (reprint of a press release)
Industrial equipmentRockwell FactoryTalk Design Studio Copilot (described as able to run on control panels and the like)9B classAnnounced (November 2025). Whether it has shipped not foundOfficial press release (13 November 2025. Described as built on Nemotron-Nano-9B-v2, fine-tuned on FactoryTalk data and run on operator panels and dedicated devices. I confirmed the body through press quotes)
GlassesGlasses with Snapdragon AR1+ Gen 11BPrototype / demo (June 2025)Press
Parts / robotsM5Stack Module LLM (fully offline)Qwen2.5-0.5B. NPU 3.2 TOPS, 4GB RAM, about 1.5WProduct (on sale)Official
Care / mental health supportResearch on a hospital robot (run locally; reason given is GDPR compliance) / research on an offline mental-health support app13B / 1BResearchPapers

For mass-produced home appliances, explaining anomalies in industrial equipment, and mass-market toys, I could not confirm from primary sources any shipped product where a small LLM runs on the device.

What I could not confirm this time

  • Muse Charm’s size, battery, chip and price (not disclosed by Meta). Screen, cameras and 5G are press reports only
  • Whether Meta’s Ray-Ban Meta glasses have an LLM on the device (I could not retrieve Meta’s official spec sheet)
  • An official list of “how many GB of RAM runs up to how many B" for phones, and Gemini Nano’s parameter count
  • An example where an official source explains an LLM running on earbuds or hearing aids
  • An example of a general-purpose small LLM of several hundred million parameters or more running on a microcontroller alone at a practical speed
  • A finished official product connecting a wearable to a local LLM at home, and official figures for Tailscale’s battery and data usage

Where would I start if putting AI in a device?

In the setup where the device only handles input and output and the brain lives at home, you need a home PC or server separate from the device. From what I found, what the reader needs to decide is where to put the AI for their own device.

What you want to buildRough guide for where the AI livesBasis
A smartphone appThe OS’s built-in on-device model (Apple, Gemini Nano), or a small model like Gemma 3nOfficial sources state the uses (summarizing, rewriting, short conversations) and supported devices
A palm-sized dedicated device (with mains power)A board in the class of Jetson Orin Nano Super: 8GB memory, 102GB/s bandwidth. The officially supported models for Raspberry Pi 5 + AI HAT+ 2 at launch include 1B-1.5B classOfficial blog measurements (Jetson), third-party measurements (Raspberry Pi)
A small device operated by voice inside the homeThe device does input and output only; the brain is a home server (Home Assistant + Ollama)An official setup. Official figures exist only for Whisper and Piper; LLM latency is a third-party reference value
A battery-powered tiny deviceI found no realistic example of running a general-purpose small LLM on a microcontroller alone. Connect to the cloud or home over a networkThe microcontroller examples are toy-sized
A device used outdoors like MuseMuse connects to the cloud. Connecting to a home server needs a combination such as a VPNMeta official (cloud). Tailscale’s official explanation

If you are unsure, the entry point with the most documentation is to install Ollama on a home PC and follow Home Assistant’s official steps (the Ollama integration) to talk to it from a small device. For ESP32-based devices, there are several products with official support.

Sources consulted

  • Meta: official Muse announcement (8 September 2026) / Meta official engineering blog “Private Processing" (23 September 2026)
  • Apple Machine Learning Research: Apple Foundation Models 2025 / the Apple Intelligence announcement of 8 June 2026
  • Google Developers Blog: Gemma 3n, Gemma 3 270M / Android Developers: Gemini Nano
  • Raspberry Pi: AI HAT+ 2 / CNX Software: AI HAT+ 2 review (20 January 2026), RK1820 and RK1828 (30 December 2025)
  • NVIDIA Developer Blog: speeding up Jetson Orin Nano / optimizing LLM inference
  • Home Assistant: Assist, the Ollama integration, official blog Voice Chapter 11 (22 October 2025)
  • Hugging Face: model cards and distribution repositories / GitHub: LLM examples on ESP32 and Pico / Arm: Small Language Model on Edge (November 2024)