Training IBM Watson using Automatically Generated Question-Answer Pairs
Lee, Kim, Yoo, Jung, Kim, Yoon · cs.CL · 2016-11-12 · 原文
IBM Watson is a cognitive computing system capable of question answering in natural languages. It is believed that IBM Watson can understand large corpora and answer relevant questions more effectively than any other question-answering system currently available. To unleash the full power of Watson, however, we need to train its instance with a large number of well-prepared question-answer pairs. Obviously, manually generating such pairs in a large quantity is prohibitively time consuming and significantly limits the efficiency of Watson's training. Recently, a large-scale dataset of over 30 million question-answer pairs was reported. Under the assumption that using such an automatically generated dataset could relieve the burden of manual question-answer generation, we tried to use this dataset to train an instance of Watson and checked the training efficiency and accuracy. According to our experiments, using this auto-generated dataset was effective for training Watson, complementing manually crafted question-answer pairs. To the best of the authors' knowledge, this work is the first attempt to use a large-scale dataset of automatically generated question-answer pairs for trainin
讲义
讲义·推断 依据「原文」自动生成的结构化摘要(推断),非原文表述;以原文为准。
1. 人话版
IBM Watson is a cognitive computing system capable of question answering in natural languages.
It is believed that IBM Watson can understand large corpora and answer relevant questions more effectively than any other question-answering system currently available.
2. 领域脉络
本文类目:cs.CL,属于其所在研究脉络的最新进展。
3. 机制拆解
摘要未展开方法细节——精读时重点看方法/模型部分。
4. 证据与数字
Recently, a large-scale dataset of over 30 million question-answer pairs was reported.
5. 反例与边界
To unleash the full power of Watson, however, we need to train its instance with a large number of well-prepared question-answer pairs.
Obviously, manually generating such pairs in a large quantity is prohibitively time consuming and significantly limits the efficiency of Watson's training.
6. 跨领域连接与意外收获
思考本文机制能否迁移到你正在跟进的问题。
7. 可复用方法
把本文机制与你手头项目对照,找一个两周内能验证的最小实验。
8. 术语表
精读时把不熟的术语记入此处,作为下次回忆的锚点。