
6/2/2026
What this post added
This post introduces Microsoft eXecution Containers (MXC) and NVIDIA OpenShell for secure, on-device agent execution on Windows PCs, enabling turnkey agent sandboxing. It details performance improvements in llama.cpp (2x for Qwen 3.5/3.6 27B, 1.6x for Qwen 3.5/3.6 35B MoE) via Multi-Token Prediction and Programmatic Dependent Launch, and in vLLM (2.6x improvement) with optimizations for MoE models and CUDA Graphs. It also highlights multi-GPU support in llama.cpp and ComfyUI through tensor parallelism, offering up to ~2x memory capacity and ~1.8x compute performance on RTX PCs. Updates to NVIDIA NemoClaw, Hermes Agent, and H Company's Holo 3.1 models are also discussed, focusing on expanded agent capabilities, ease of setup, native app integration, and performance on NVIDIA GPUs.