Constarium
← Search

Text · dataset · 2022

transformersbook/codeparrot

Listed in Hugging Face Datasets

CodeParrot 🦜 Dataset What is it?

Description

This is the full CodeParrot dataset. It contains Python files used to train the code generation model in Chapter 10: Training Transformers from Scratch in the NLP with Transformers book.

You can find the full code in the accompanying Github repository. Creation It was created with the GitHub dataset available via Google's BigQuery. It contains approximately 22 million Python files and is 180 GB (50 GB compressed) big.

Read the rest (1 more)

The… See the full description on the dataset page: huggingface.co/datasets/transformersbook/codeparrot.

Links

Where it is published

Catalogue records · 1

Topics

Stated by source
text
Provenance · 1 source records, 8 field assertions
SourceKeyLast seenRaw
Hugging Face Datasetstransformersbook/codeparrot12 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:textsource · Hugging Faceconnector:huggingface@1.0.0
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0