数据洋

wikiquote-en-simplified

Movies and TV ShowsComputer Science

￥30

13.59MB

数据标识：D17168903356580243

发布时间：2024/05/28

About Dataset

Data Collection

The data is parsed using ceshine/wickedQuotes, which is a fork of heyseth / wickedQuotes by Seth Miller.

I did some refactoring to the original code and convert the JSON dump into two CSV files (one containing source information, while the other containing the quotes).

Note that Seth Miller also published a Wikiquote dataset here on Kaggle, but it is in JSON format.

The 2020-01-20 Wikiquote dump is used to generate this dataset.

Additional File information

We provide two set of quotes:

The quotes-100-en-* files contain quotes that are shorter than 100 characters.
The quotes-500-en-* files contain quotes that are shorter than 500 characters.

看了又看

验证报告

当前版本暂不支持对此种交付方式或数据格式开展数据质量验证，相关校验能力将在后续版本上线，敬请期待。

wikiquote-en-simplified

￥30

13.59MB

申请报告

wikiquote-en-simplified

About Dataset

Data Collection

Additional File information

关于典枢

下载与支持

服务协议

关于我们

官方公众号

技术交流群