Skip to main content
NeuroActu — Your daily watch on artificial intelligence
NeuroActu AI news, every day
Models Local LLM AI Agents AI Creation Hardware Tools Companies Regulation
Category · 9 articles

Local LLM

Open-source models to run on your own machine: Llama, Mistral, Qwen, Gemma, Phi, DeepSeek… Quantization, fine-tuning, hardware optimizations, benchmarks and inference tools (llama.cpp, ollama, vLLM).

Kimi 3 aims to challenge Opus 4.8 with a Chinese open giant Local LLM
17 juillet 2026 ·TechCrunch AI · 4 min

Kimi 3 aims to challenge Opus 4.8 with a Chinese open giant

Moonshot is reportedly preparing Kimi 3, a Chinese open model with 2 to 3 trillion parameters, designed to rival Claude Opus 4.8.

Read the article
Hugging Face wants to speed up LLMs with diffusion Local LLM
23 mai 2026 ·Hugging Face · 4 min

Hugging Face wants to speed up LLMs with diffusion

Hugging Face is highlighting diffusion language models with Nemotron-Labs, a promising path to dramatically speed up text generation.

Read the article
Hugging Face speeds up LLMs with Nemotron-Labs Local LLM
23 mai 2026 ·Hugging Face · 4 min

Hugging Face speeds up LLMs with Nemotron-Labs

Hugging Face is highlighting Nemotron-Labs’ diffusion language models, a promising path to significantly speed up text generation.

Read the article
Hugging Face wants near-instant LLMs through diffusion Local LLM
25 mai 2026 ·Hugging Face · 4 min

Hugging Face wants near-instant LLMs through diffusion

Hugging Face is highlighting Nemotron-Labs’ diffusion language models to greatly speed up text generation and challenge autoregressive LLMs.

Read the article
Open-source LLMs locally: which ones run on a PC in 2026? Local LLM
14 juillet 2026 ·NeuroActu · 22 min

Open-source LLMs locally: which ones run on a PC in 2026?

2026 editorial comparison of open-source LLMs that are genuinely usable locally on a standard PC: RAM, VRAM, speed, quality, privacy, and trade-offs.

Read the article
DiffusionGemma: Google promises text 4x faster Local LLM
11 juin 2026 ·Google DeepMind · 4 min

DiffusionGemma: Google promises text 4x faster

Google DeepMind introduces DiffusionGemma, an approach that promises text generation up to 4x faster than traditional LLMs.

Read the article
GLM-5.2 aims to make open source matter for long tasks Local LLM
17 juin 2026 ·Hugging Face · 4 min

GLM-5.2 aims to make open source matter for long tasks

Hugging Face introduces GLM-5.2, a model designed for long tasks and persistent agents, with an open-source positioning worth watching.

Read the article
Hugging Face simplifies one-click vLLM deployment Local LLM
28 juin 2026 ·Hugging Face · 4 min

Hugging Face simplifies one-click vLLM deployment

Hugging Face now makes it possible to launch a vLLM server with a single command via HF Jobs, a concrete step toward industrializing open source LLMs.

Read the article
GLM-5.2 aims to rival Anthropic in cybersecurity Local LLM
29 juin 2026 ·The Verge AI · 4 min

GLM-5.2 aims to rival Anthropic in cybersecurity

Z.ai unveils GLM-5.2, an open-weight model presented as capable of rivaling Mythos on certain cybersecurity tasks.

Read the article