DS数据代找

wikiquote-en-simplified

Movies and TV ShowsComputer Science

30

已售 0
13.59MB

数据标识:D17168903356580243

发布时间:2024/05/28

卖家暂未授权典枢平台对该文件进行数据验证,您可以向卖家

申请验证报告

数据描述

About Dataset

Data Collection

The data is parsed using ceshine/wickedQuotes, which is a fork of heyseth / wickedQuotes by Seth Miller.

I did some refactoring to the original code and convert the JSON dump into two CSV files (one containing source information, while the other containing the quotes).

Note that Seth Miller also published a Wikiquote dataset here on Kaggle, but it is in JSON format.

The 2020-01-20 Wikiquote dump is used to generate this dataset.

Additional File information

We provide two set of quotes:

  • The quotes-100-en-* files contain quotes that are shorter than 100 characters.
  • The quotes-500-en-* files contain quotes that are shorter than 500 characters.
data icon
wikiquote-en-simplified
30
已售 0
13.59MB
申请报告