Down Shift

Stanford Question Answering Dataset (SQuAD)

educationnlptext miningtexttext pre-processing

￥5

10.08MB

数据标识：D17171515568624732

发布时间：2024/05/31

The Stanford Question Answering Dataset

A Challenge for Reading Comprehension

About this dataset

> SQuAD is a reading comprehension dataset consisting of questions posed by crowdworkers on a set of Wikipedia articles. The answers to the questions are span of text, or segments, from the corresponding reading passages. The data fields in this dataset are the same across all splits

How to use the dataset

> The SQuAD dataset is a reading comprehension dataset, consisting of questions posed by crowdworkers on a set of Wikipedia articles, where the answer to every question is a segment of text, or span, from the corresponding reading passage. The data fields are the same among all splits > > Columns:context,question,answers > > To use this dataset, simply download one of the split files (train.csv or validation.csv) and load it into your preferred data analysis tool. Each row in the file corresponds to a single question-answer pair. The context column contains the full text of the corresponding Wikipedia article, while the question and answers columns contain the question posed by the crowdworker and its corresponding answer(s)

Research Ideas

> - Learning to answer multiple choice questions by extracting text spans from source materials > - Developing Reading Comprehension models that can answer open-ended questions about passages of text > - Building systems that can generate large training datasets for Reading Comprehension models by creating synthetic questions from existing passages

Acknowledgements

> Thank you to the Stanford Natural Language Inference group and the creators of the SQuAD dataset for providing this data

> > > ### License > > > > License: CC0 1.0 Universal (CC0 1.0) - Public Domain Dedication > > No Copyright - You can copy, modify, distribute and perform the work, even for commercial purposes, all without asking permission. See Other Information.

Columns

File: validation.csv

Column name	Description
title	The title of the Wikipedia article. (String)
context	The full text of the article. (String)
question	The question posed by the crowdworker. (String)
answers	The answer to the question, as a string of text spans. (List of strings)

File: train.csv

Column name	Description
title	The title of the Wikipedia article. (String)
context	The full text of the article. (String)
question	The question posed by the crowdworker. (String)
answers	The answer to the question, as a string of text spans. (List of strings)

看了又看

验证报告

以下为卖家选择提供的数据验证报告：

Stanford Question Answering Dataset (SQuAD)

￥5

10.08MB

申请报告

Stanford Question Answering Dataset (SQuAD)

The Stanford Question Answering Dataset

A Challenge for Reading Comprehension

About this dataset

How to use the dataset

Research Ideas

Acknowledgements

Columns

关于典枢

下载与支持

服务协议

关于我们

官方公众号

技术交流群