ถามวิดีโอด้วยภาษาคน: Agentic AI ตอบภายในวินาทีบน AWS

แนะนำสถาปัตยกรรม “Video Intelligence แบบ Agentic” ค้นหาความรู้ในวิดีโอได้ทันที ลดข้อจำกัดการเข้าถึงการเรียนรู้ พร้อมซอร์สโค้ดตัวอย่าง
เมื่อต้องไล่ดูคลิปประชุมยาวหลายชั่วโมงหรือฟุตเทจความปลอดภัยนานนับสัปดาห์ การเข้าถึง “ช่วงเวลาสำคัญ” กลายเป็นอุปสรรคใหญ่ต่อการเรียนรู้และการตัดสินใจ ทีมเทคโนโลยีจำนวนมากรวมถึงนักศึกษาใช้เวลามากเกินไปกับการกรอวิดีโอแทนการวิเคราะห์สาระสำคัญ SPU ชี้แนวทางแก้ปัญหานี้ผ่านสถาปัตยกรรม Video Intelligence แบบ Agentic ที่ให้ผู้ใช้ถามคำถามด้วยภาษาธรรมชาติและรับคำตอบจากเนื้อหาในวิดีโอได้ภายในไม่กี่วินาที
แนวทางนี้ใช้ Strands Agents SDK สร้าง “เอเจนต์เดียว” ที่ตัดสินใจเอง ณ เวลาใช้งานว่าจะเรียกใช้บริการใดของ AWS ตามคำถามจริง ไม่ต้องสร้างไปป์ไลน์เฉพาะกิจสำหรับแต่ละยูสเคสอีกต่อไป เอเจนต์จะ:
– ใช้ Amazon Transcribe เมื่อต้องวิเคราะห์เสียงพูด
– ใช้ Amazon Rekognition เมื่อต้องค้นหาภาพ ใบหน้า ฉาก หรือกิจกรรม
– เรียก Amazon Bedrock Data Automation (BDA) เมื่อผู้ใช้ต้องการสรุปรวมทั้งคลิปในครั้งเดียว
– ใช้ผลการวิเคราะห์ที่แคชไว้สำหรับคลิปเดิมเพื่อตอบคำถามต่อเนื่องในเวลาไม่ถึงหนึ่งวินาที
ผลการทดสอบตัวอย่างกับวิดีโอ 60–90 นาที พบว่า คำถามแรกของคลิปใหม่ใช้เวลา 5–10 นาทีระหว่างที่ระบบถอดเสียงหรือวิเคราะห์ภาพ ส่วนคำถามถัดไปตอบแทบจะทันทีเพราะใช้แคช ที่สำคัญ ผู้ใช้สามารถถามเชิงสืบสวนได้ เช่น “ช่วงไหนที่รถเปลี่ยนเลนก่อนชน” หรือคำถามด้านเนื้อหา “ในที่ประชุมนี้มีการตัดสินใจเรื่องใดบ้าง” โดยระบบสังเคราะห์คำตอบพร้อมไทม์สแตมป์อ้างอิง
ตัวอย่างอินเทอร์เฟซเป็นแชตพร้อมอัปโหลดไฟล์และเลือกโหมดวิเคราะห์ได้ตามต้องการ (ดูภาพ Figure 1 ในต้นฉบับ) โครงสร้างระบบ (Figure 2) แยกบทบาทชัดเจน: Agent Orchestrator ใช้ Amazon Bedrock เป็นสมอง, Amazon S3 เก็บไฟล์และผลวิเคราะห์แบบแยกผู้ใช้, และเพิ่ม Amazon Bedrock Guardrails ได้เมื่อต้องการคุมเนื้อหาที่อ่อนไหว
การดีพลอยอ้างอิงใช้ Amazon ECS กับ AWS Fargate วางหลัง Amazon CloudFront ที่เป็นจุดเข้าเพียงจุดเดียว ส่วนการยืนยันตัวตนใช้ Amazon Cognito แบบเชิญเท่านั้นและบังคับ MFA ข้อมูลถูกจัดเก็บใน S3 ภายใต้พรีฟิกซ์ต่อผู้ใช้ พร้อมนโยบายลบอัตโนมัติ 24 ชั่วโมง (ดู Figure 4–5)
เคสการใช้งานครอบคลุมตั้งแต่ Meeting Intelligence, การเฝ้าระวังการเข้า-ออกอาคาร จนถึงงานสืบสวนเคลมประกัน โดยมีองค์กรสื่อรายใหญ่ซึ่งรับคำปรึกษา AWS Professional Services รายงานว่าลดเวลาทบทวนแบบแมนนวลได้ราว 80% จากแบ็กล็อกมากกว่า 200 คลิปหลายชั่วโมง (เป็นข้อมูลรายงานภายในของลูกค้า ไม่ได้ตรวจสอบโดยอิสระ)
เพื่อการเรียนรู้เชิงปฏิบัติ เอเจนต์ โค้ด และสคริปต์ดีพลอยถูกเปิดเป็นทรัพยากรสาธารณะ:
– บทความและสถาปัตยกรรมตัวอย่าง: https://github.com/aws-samples/sample-media-analysis-agent
– บริการที่เกี่ยวข้อง: Amazon Bedrock, Amazon Rekognition, Amazon Transcribe, Strands Agents SDK
สำหรับนักศึกษาและผู้เชี่ยวชาญด้านเทคโนโลยี โซลูชันนี้ลดข้อจำกัดการเข้าถึงความรู้ที่ “ซ่อนอยู่” ในวิดีโอจำนวนมหาศาล ช่วยเปลี่ยนชั่วโมงการไล่ดูเป็นคำตอบที่ค้นหาได้ด้วยภาษาคน และขยายการเรียนรู้ตลอดชีวิตด้วยเครื่องมือเปิดที่พร้อมทดลองใช้งานทันที

Spotlights an agentic video intelligence blueprint that unlocks learning from long-form videos—open-source code included
When hours of recordings bury the moments that matter, teams—and students—spend too much time scrubbing timelines instead of learning. SPU highlights a practical agentic architecture that turns videos into a conversational knowledge source: ask a natural question, get an answer with timestamps in seconds.
Built with the Strands Agents SDK, the solution orchestrates AWS services at runtime based on the user’s query—no per-use-case pipelines. The agent:
– Invokes Amazon Transcribe for spoken-content queries
– Uses Amazon Rekognition for visual search and face matching
– Calls Amazon Bedrock Data Automation (BDA) for comprehensive “summarize this video” requests
– Reuses cached results so follow-up questions on the same video return in under a second
In tests with 60–90 minute videos, the first question typically takes 5–10 minutes while transcription or visual analysis runs; subsequent queries are near-instant thanks to caching. The chat UI supports uploads and analysis modes, and the architecture cleanly separates roles: Bedrock as the reasoning engine, Amazon S3 for per-user storage and cached results, with optional Bedrock Guardrails for responsible-AI controls.
The reference deployment runs on Amazon ECS with AWS Fargate behind Amazon CloudFront, and uses Amazon Cognito (invitation-only with MFA) for secure access. Per-user S3 prefixes and quotas provide multi-tenant isolation. Example use cases span meeting intelligence, access monitoring, and insurance investigations. One media and entertainment company working with AWS Professional Services reported about an 80% reduction in manual review time across 200+ multi-hour recordings (customer-reported; not independently verified).
For hands-on learning, the complete implementation and deployment scripts are public:
– Source and reference architecture: https://github.com/aws-samples/sample-media-analysis-agent
– Related services: Amazon Bedrock, Amazon Rekognition, Amazon Transcribe, Strands Agents SDK
For SPU’s tech community, this pattern lowers the barrier to learning from large video backlogs—converting time-consuming reviews into fast, natural-language answers and supporting lifelong learning with open resources you can try today.

ที่มา: https://aws.amazon.com/blogs/machine-learning/agentic-conversational-video-intelligence-built-on-aws/

Most Popular

Categories