Skip to content
NexLMAI
Hugging Face ↗
Independent artificial intelligence
Built for the edge
Building now — 2026

Intelligence,
made closer.

Capable language models that keep powerful intelligence in your hands—not somewhere else.

Explore models → Built for local inference ↓
03 / Local by design

Your model.
Your machine.

NexLM models are built for on-device inference: private, responsive, and available without a round trip to the cloud.

  1. 01Choose a runtime below.
  2. 02Copy the fetch command for your stack.
  3. 03Run privately, on your hardware.
nexlm — local / grm-3-nano
$ nexlm
  • ▶Run a model
  • ▶Launch with Ollama
  • ▶Launch with MLX
  • ▶Launch with llama.cppnot installed
  • ▶More…
Fetch command
ollama pull nexlm/grm-3-nano
NexLM / Model index

Models

A focused family of models for reasoning, local deployment, code, and specialized work. Select a system to view its current direction.

Six systems, one direction01—06
Status Released Finishing eval In development Security / approval gated
NexLM / Hub

Models, releases,
and artifacts.

Find NexLM model pages and published releases on Hugging Face.

Visit NexLM on Hugging Face ↗
NexLM / Dispatches

News

Product milestones, research progress, and notes from a company building intelligence for the device in front of you.

Latest dispatches2026

NexLM / Dispatch Return to news ↑
NexLM / Research

Research

NexLM research is centered on capable, efficient language models that can move beyond the data center. We are documenting the methods, systems work, and evaluations behind each release as they mature.

01 / Architecture

Efficient capability

We investigate architectures that preserve useful reasoning and general performance while remaining practical to deploy close to the user.

02 / Post-training

Behavior with intent

We refine model behavior for useful instruction following, technical work, and domain-specific tasks without losing directness.

03 / Evaluation

Real-world signals

Beyond headline benchmarks, we examine coding workflows, structured reasoning, local latency, and hands-on model behavior.

04 / Deployment

Inference that travels

Our deployment work considers MLX, GGUF, Ollama, and llama.cpp so models can operate across developer and edge environments.

NexLM / Contact

Make intelligence yours.

For model access, research conversations, deployment questions, or collaboration.

nexlm@icloud.com ↗

© 2026 NexLM AI Hugging Face ↗ Made with intent