ADP 前沿学习

← 板块一 · 研究前沿

Training IBM Watson using Automatically Generated Question-Answer Pairs

Lee, Kim, Yoo, Jung, Kim, Yoon · cs.CL · 2016-11-12 · 原文

IBM Watson is a cognitive computing system capable of question answering in natural languages. It is believed that IBM Watson can understand large corpora and answer relevant questions more effectively than any other question-answering system currently available. To unleash the full power of Watson, however, we need to train its instance with a large number of well-prepared question-answer pairs. Obviously, manually generating such pairs in a large quantity is prohibitively time consuming and significantly limits the efficiency of Watson's training. Recently, a large-scale dataset of over 30 million question-answer pairs was reported. Under the assumption that using such an automatically generated dataset could relieve the burden of manual question-answer generation, we tried to use this dataset to train an instance of Watson and checked the training efficiency and accuracy. According to our experiments, using this auto-generated dataset was effective for training Watson, complementing manually crafted question-answer pairs. To the best of the authors' knowledge, this work is the first attempt to use a large-scale dataset of automatically generated question-answer pairs for trainin

🔮 让 ChatGPT 全网深度追问

讲义

讲义·推断 依据「原文」自动生成的结构化摘要(推断),非原文表述;以原文为准。

1. 人话版

IBM Watson is a cognitive computing system capable of question answering in natural languages.

It is believed that IBM Watson can understand large corpora and answer relevant questions more effectively than any other question-answering system currently available.

2. 领域脉络

本文类目:cs.CL,属于其所在研究脉络的最新进展。

3. 机制拆解

摘要未展开方法细节——精读时重点看方法/模型部分。

4. 证据与数字

Recently, a large-scale dataset of over 30 million question-answer pairs was reported.

5. 反例与边界

To unleash the full power of Watson, however, we need to train its instance with a large number of well-prepared question-answer pairs.

Obviously, manually generating such pairs in a large quantity is prohibitively time consuming and significantly limits the efficiency of Watson's training.

6. 跨领域连接与意外收获

思考本文机制能否迁移到你正在跟进的问题。

7. 可复用方法

把本文机制与你手头项目对照,找一个两周内能验证的最小实验。

8. 术语表

精读时把不熟的术语记入此处,作为下次回忆的锚点。