PAIR: Router เสมือนที่กระจายงาน LLM ข้ามเครื่องในบ้านได้ทันที—ไม่ต้องแก้โค้ด ไม่ต้องตั้งคลัสเตอร์

แก้คอขวดมัลติเอเจนต์บนเครื่องเดียว ให้โน้ตบุ๊ก พีซี และ DGX Spark ทำงานพร้อมกัน—NVIDIA Personal AI Router เปิดทางให้นักศึกษาทดลอง AI ที่บ้านได้เร็วขึ้น
เมื่อเวิร์กโฟลว์แบบมัลติเอเจนต์กลายเป็นมาตรฐาน งานหนึ่งคำขอกลายเป็นหลายสิบคำสั่งย่อยที่ยิงเข้าเอนจินเดียวพร้อมกัน คิวหนาแน่น ขณะเดียวกันเครื่องแรงๆ ในบ้านหรือในเครือข่ายเดียวกันกลับว่างงาน นี่คือคอขวดที่ NVIDIA Personal AI Router (PAIR) ตั้งใจปลดล็อก
ประกาศสัปดาห์นี้ PAIR คือ “virtual inference router” ที่ค้นหาเครื่องที่รองรับบนเครือข่ายภายในบ้าน แล้วกระจายคำขอ inference อิสระไปยังเครื่องเหล่านั้นโดยอัตโนมัติ ไม่ใช่เอนจินใหม่—Ollama หรือ LM Studio ยังเป็นตัวรันโมเดลบนโหนดที่ PAIR เลือกให้ จุดเด่นคือไม่ต้องเปลี่ยนฮาร์เนสเอเจนต์เดิม เพราะ PAIR ทำตัวเป็นพร็อกซีเข้ากับอินเทอร์เฟซที่ Ollama/LM Studio ใช้อยู่แล้ว และยังมีพร็อกซีแบบ OpenAI-compatible ให้ด้วย
การใช้งานทำได้ทันที—PAIR เปิดให้ทดสอบสาธารณะ (v0.1.1) พร้อมตัวติดตั้งที่ลงนามอย่างเป็นทางการบน Windows, macOS, และ Linux โค้ดทั้งหมดเปิดที่ GitHub ภายใต้ Apache 2.0 ทำงานบนเครือข่ายท้องถิ่น โดยต้องใช้อินเทอร์เน็ตเฉพาะตอนดาวน์โหลดโมเดล
การค้นหาและจับคู่โหนดทำผ่าน mDNS หากค้นหาไม่เจอสามารถเพิ่มด้วย IP ได้ ความเชื่อใจเริ่มจาก PIN 6 หลักที่แสดงบนเครื่องเชิญและต้องกรอกบนเครื่องที่ถูกเชิญ การสื่อสารระหว่างโหนดจะถูกบล็อกจนกว่าจะจับคู่สำเร็จ จากนั้นทราฟฟิกจะถูกเข้ารหัสด้วย mTLS พร้อมใบรับรองที่สร้างอัตโนมัติ แต่ละโหนดต้องมี Ollama หรือ LM Studio—PAIR ช่วยติดตั้งเอนจินและสั่งดาวน์โหลดโมเดลข้ามเครื่องให้ได้ ลดงานเซ็ตอัปแบบหลายเครื่องไปได้มาก
ตัวจัดตารางของ PAIR จะพิจารณาโหนดที่ “พร้อมจริง” เท่านั้น: เอนจินที่รองรับถูกเปิดใช้งานอยู่, มีโมเดล “ตรงชื่อและแท็ก” ตามที่ร้องขอ, ภาระงานปัจจุบันของโหนด/เอนจิน และการใช้งาน GPU ณ ขณะนั้น โมเดลไม่จำเป็นต้องเหมือนกันทุกเครื่อง—ใครมีอะไร PAIR จะส่งงานไปตามนั้น ถ้าโหลดแท็กเดียวกันไว้หลายเครื่อง ก็เพียงเพิ่มกลุ่มโหนดที่เข้าข่ายเลือกได้ นี่คือ concurrency ระดับเวิร์กโหลดอย่างชัดเจน: คำขอหนึ่งคำขอจะถูกผูกกับโหนดเดียวตลอดอายุงาน ไม่รวม VRAM ไม่รวม GPU คนละเครื่องให้เป็นตัวเร่งใหญ่ตัวเดียว และไม่แตกงานหนึ่งคำขอไปข้ามเครื่อง
เดโมของ NVIDIA จับคู่ PAIR เข้ากับ Hermes Desktop สร้างเวิร์กโหลด 5 ซับเอเจนต์บนกล่องอีเมลจำลอง โดย Ollama รัน Qwen 3.6 35B A3B บนโหนดที่ถูกเลือก ผลคือบนแล็ปท็อป RTX Spark เครื่องเดียว ใช้เวลาเฉลี่ย 18 นาที แต่เมื่อรันบนคลัสเตอร์ 3 เครื่อง (RTX Spark laptop, DGX Spark และ RTX 5090) เวลาลดลงเหลือเฉลี่ย 8 นาที 48 วินาที
ฮาร์ดแวร์ที่รองรับครอบคลุม GeForce RTX 20 Series ขึ้นไป, RTX PRO ตั้งแต่ Turing, DGX Spark และ Apple Silicon รุ่น M4 ขึ้นไป จับคู่ข้าม Windows, Linux, macOS ได้ทั้ง x64 และ arm64 (Windows on ARM ยังเป็น experimental) การคอนฟิกที่ยืนยันแล้วแนะนำ RAM 8 GB ขึ้นไป และพื้นที่ดิสก์อย่างน้อย 20 GB ดิสโทร Linux อื่นๆ สามารถบิลด์จากซอร์ส
สาระสำคัญสำหรับสายเทค:
– PAIR กระจาย “คำขออิสระ” ไปยังโหนดเดียวต่อคำขอ—ไม่รวม VRAM/ไม่ชาร์ดงานเดียวข้ามเครื่อง
– พร็อกซีอินเทอร์เฟซของ Ollama/LM Studio เดิม—ฮาร์เนสเอเจนต์ไม่ต้องแก้โค้ด
– เกณฑ์คัดเลือกโหนด: ความพร้อม, เอนจินที่รองรับ, โมเดลตรงแท็ก, ภาระงาน, การใช้ GPU
– เดโม 5 ซับเอเจนต์ ลดเวลาเฉลี่ยจาก 18 นาที เหลือ 8:48 บน 3 อุปกรณ์ (ตัวเลขไม่เป็นทางการ)
– สถานะเบตา Apache 2.0—นโยบายจัดตารางปัจจุบันยังไม่คำนึงถึงขนาด VRAM คลาส GPU หรือความอุ่นของโมเดล
สำหรับนักศึกษา นักพัฒนา และคณาจารย์ที่ต้องทดสอบเวิร์กโฟลว์มัลติเอเจนต์หรือ LLM บนแล็บส่วนตัว PAIR เปิดทางให้รีดศักยภาพเครื่องว่างในบ้านหรือออฟฟิศย่อยได้ทันที โดยไม่ต้องผูกมัดกับโค้ดใหม่หรือสถาปัตยกรรมคลัสเตอร์ซับซ้อน ช่วยลดข้อจำกัดในการเข้าถึงการเรียนรู้เชิงปฏิบัติด้าน AI อย่างเป็นรูปธรรม

Goodbye single-node bottlenecks—NVIDIA Personal AI Router spreads multi‑agent workloads across laptops, PCs, and DGX Spark so students can learn and experiment faster
Multi-agent workflows turned one prompt into dozens of parallel model calls—hammering a single local engine while other capable machines sit idle on the same network. NVIDIA’s Personal AI Router (PAIR), announced this week, targets that exact choke point.
PAIR is a virtual inference router that discovers compatible systems on your local network and schedules independent requests across them. It is not a new inference engine—Ollama or LM Studio still executes the model on whichever node PAIR selects. Critically, PAIR introduces no new cluster API. It proxies the Ollama-compatible and LM Studio-compatible endpoints (and exposes OpenAI-compatible proxies), taking over the engines’ default ports or a configurable proxy port. Your agent harness stays unchanged: the agent decides the work; PAIR decides where it runs.
Deployability is straightforward: PAIR ships as a public beta (v0.1.1) with signed installers for Windows, macOS, and Linux, and full source on GitHub under Apache 2.0. It runs entirely on the local network; the internet is needed only to download models.
Discovery and pairing use mDNS, with manual IP add when needed. Trust is established via a six-digit PIN shown on the inviter and entered on the invitee; all traffic is blocked until pairing completes. Communication between paired nodes is secured with mTLS using generated certificates. Each node runs Ollama or LM Studio, and PAIR can help install engines and trigger model downloads across machines to reduce multi-device setup friction.
Scheduling respects explicit workload boundaries. A node is eligible only when a supported engine is enabled and the exact requested model tag is present. Models can differ across machines; routing follows model location, and loading the same tag on more nodes widens the eligible pool. For each request PAIR weighs five signals: node readiness, engine availability, exact model presence, current node/engine job load, and current GPU utilization. PAIR never pools VRAM, merges GPUs, or shards a single request across machines.
Demo numbers (with caveats): paired with Hermes Desktop creating a five‑subagent workload over a synthetic inbox, Ollama runs Qwen 3.6 35B A3B on each selected node. One RTX Spark laptop averaged 18 minutes. A three‑device cluster (RTX Spark laptop, DGX Spark, RTX 5090) averaged 8:48.
Supported platforms include GeForce RTX 20 Series and newer, RTX PRO (Turing onward), DGX Spark, and Apple M4 or newer silicon, across Windows, Linux, and macOS on x64 and arm64 (Windows on ARM is experimental). Validated configs list 8 GB RAM+ and ~20 GB recommended disk; other Linux distributions build from source.
Key takeaways for builders:
– One independent request per node; no VRAM pooling and no model sharding
– Proxies existing Ollama/LM Studio endpoints, so harnesses need no changes
– Eligibility: readiness, engine, exact model, job load, and GPU utilization
– Unofficial demo: 18 min on one laptop vs 8:48 on three devices (five subagents)
– Apache 2.0 beta with a single scheduling policy—currently blind to VRAM, GPU class, and model warmness
For students, developers, and faculty running multi‑agent workloads in home labs, PAIR offers a practical path to reclaim idle GPUs and accelerate hands‑on LLM experimentation—reducing barriers to learning without adding cluster complexity.

ที่มา: https://www.marktechpost.com/2026/09/04/nvidia-releases-personal-ai-router-pair-an-open-source-virtual-inference-router-that-distributes-local-ai-requests-across-rtx-dgx-spark-and-mac-nodes/

Most Popular

Categories