数据洋

writingqualitymemoryreduction

LinguisticsRegressionData CleaningTabular

￥20

74MB

数据标识：D17168924748077832

发布时间：2024/05/28

About Dataset

This is a memory reduced dataset for the Writing Process-Writing Quality competition. I encoded text columns into np.int8 type and binned categories with extremely low occurrences into a common bin. I also down-casted certain columns in the data based on their min-max values to save memory. I have saved the train-logs data in a binary format and the encoded text strings and their categories too as one may need them while inferring on the test data.
This is also available in my baseline data prep kernel.
We will use this data as input for all our future steps including EDA, model development and inference development. We hope not to fall prey to memory errors using such an approach.
All the best for the competition!

看了又看

验证报告

当前版本暂不支持对此种交付方式或数据格式开展数据质量验证，相关校验能力将在后续版本上线，敬请期待。

writingqualitymemoryreduction

￥20

74MB

申请报告

writingqualitymemoryreduction

About Dataset

关于典枢

下载与支持

服务协议

关于我们

官方公众号

技术交流群