About Dataset
Data Collection
The data is parsed using ceshine/wickedQuotes, which is a fork of heyseth / wickedQuotes by Seth Miller.
I did some refactoring to the original code and convert the JSON dump into two CSV files (one containing source information, while the other containing the quotes).
Note that Seth Miller also published a Wikiquote dataset here on Kaggle, but it is in JSON format.
The 2020-01-20 Wikiquote dump is used to generate this dataset.
Additional File information
We provide two set of quotes:
- The
quotes-100-en-*files contain quotes that are shorter than 100 characters. - The
quotes-500-en-*files contain quotes that are shorter than 500 characters.
看了又看
暂无推荐
验证报告

目前该文件尚无匹配的数据质量验证程序。我们将在后续版本中提供相应的验证支持,敬请谅解。

wikiquote-en-simplified
13.59MB
申请报告




