Preprocessed Python Code Corpus
Published in Zenodo (CERN European Organization for Nuclear Research) • Jan 27, 2020
NobleIDNI8P16W46R41S93
Authors:,,
Rafael-Michael Karampatsis
Hlib Babii
Romain Robbes
Abstract
A preprocessed code corpus for the Python programming language. The corpus was used for the experiments in the paper Big Code != Big Vocabulary: Open-Vocabulary Models for Source Code. It contains preprocessed-tokenized files for training, validation, testing, and BPE encoding learning. The BPE segm...
Finding related papers...
Discussions
(0)No comments yet
Be the first to share your thoughts!