本記事は SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models (arXiv:2303.08896) の解説記事です。 論文概要(Abstract) SelfCheckGPTは、外部知識ベースを必要とせずにLLMの幻覚(hallucination...
24/03/2026 blog paper
LLM hallucination detection +5
本記事は Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (arXiv:2306.05685) の解説記事です。 論文概要(Abstract) 本論文は、LLMを評価の判定器として使う「LLM-as-a-Judge」パラダイムの信頼性を体系的に検証した研究である。著者らは、多段階対話を評価するMT-Benchと、人間のペアワイ...
24/03/2026 blog paper
LLM evaluation LLM-as-a-Judge +4
本記事は Benchmarking Large Language Models in Retrieval-Augmented Generation (arXiv:2312.10997) の解説記事です。 論文概要(Abstract) RGB(Retrieval-Augmented Generation Benchmark)は、RAGシステムに必要な4つの基本能力を体系的に評価するベンチマ...
24/03/2026 blog paper
RAG benchmark evaluation +4
本記事は RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models (ACL 2024) の解説記事です。 論文概要(Abstract) RAGTruthは、RAG(Retrieval-Augmented Generation)システムが生成する幻覚(ha...
24/03/2026 blog paper
RAG hallucination evaluation +4
本記事は RAGAS: Automated Evaluation of Retrieval Augmented Generation (arXiv:2309.01431) の解説記事です。 論文概要(Abstract) RAGASは、Retrieval-Augmented Generation(RAG)パイプラインを参照回答なし(reference-free)で自動評価するフレームワーク...
24/03/2026 blog paper
RAG evaluation LLM +4
本記事は Code execution with MCP: building more efficient AI agents (Anthropic Engineering Blog) の解説記事です。 ブログ概要(Summary) Anthropicのエンジニアリングブログは、MCP(Model Context Protocol)を通じたツール連携におけるトークン消費の非効率性を指摘し...
23/03/2026 blog tech_blog
MCP code-execution token-optimization +3
本記事は Adaptive Orchestration: Cognitive Architectures in Multi-Agent AI Systems for Enterprise Applications (arXiv:2502.18082) の解説記事です。 論文概要(Abstract) 本論文は、エンタープライズ向けマルチエージェントAIシステムのための認知アーキテクチャ「A...
23/03/2026 blog paper
multi-agent orchestration circuit-breaker +4
本記事は FlowBench: Evaluating LLM Agents Across Diverse Procedural Workflows (arXiv:2503.12347) の解説記事です。 論文概要(Abstract) FlowBenchは、LLMエージェントを4つのワークフロー型(Sequential、Conditional、Parallel、Recovery)で体系的に...
23/03/2026 blog paper
LLM agent benchmark +4
本記事は Tool-Augmented LLMs: A Survey on Integration Architectures and Failure Patterns (arXiv:2504.10376) の解説記事です。 論文概要(Abstract) 本論文は、ツール拡張LLM(Tool-Augmented LLMs)における統合アーキテクチャと障害パターンを体系的に整理したサーベイ...
23/03/2026 blog paper
LLM tool-use failure-patterns +4
本記事は MCP-Zero: Active Tool Discovery and Recommendation for LLM Agents (arXiv:2504.08999) の解説記事です。 論文概要(Abstract) MCP-Zero は、300以上のMCP(Model Context Protocol)ツールが存在する大規模環境において、LLMエージェントがタスクに必要なツー...
23/03/2026 blog paper
MCP tool-discovery circuit-breaker +3