Hugging Face Datasets2023 · dataset · gated
Dialog StudioDialogStudio: Unified Dialog Datasets and Instruction-Aware Models for Conversational AI Author: Jianguo Zhang, Kun Qian Paper|Github|[GDrive] 🎉 March 18, 2024: Update for AI Agent. Check xLAM for the latest data and models relevant to AI Agent! 🎉 March 10 2024: Update for dataset viewer issues: Please refer to https://github.com/salesforce/DialogStudio for view of each dataset, where we provide 5
Hugging Face Datasets2022 · Text
XAlignIt consists of an extensive collection of a high quality cross-lingual fact-to-text dataset where facts are in English and corresponding sentences are in native language for person biographies. The Train & validation splits are created using distant supervision methods and Test data is generated through human annotations.
Hugging Face Datasets2022 · Text
TaTADataset loader for TaTA: A Multilingual Table-to-Text Dataset for African Languages
Hugging Face Datasets2022 · dataset
RUCAIBox/Data-to-text-GenerationThis is the data-to-text generation datasets collected by TextBox, including: WebNLG v2.1 (webnlg) WebNLG v3.0 (webnlg2) WikiBio (wikibio) E2E (e2e) DART (dart) ToTTo (totto) ENT-DESC (ent) AGENDA (agenda) GenWiki (genwiki) TEKGEN (tekgen) LogicNLG (logicnlg) WikiTableT (wikit) WEATHERGOV (wg). The detail and leaderboard of each dataset can be found in TextBox page.
Hugging Face Datasets2022 · Text
mlb_data_to_textThe MLB dataset for data to text generation contains Major League Baseball games statistics and their human-written summaries.
Hugging Face Datasets2022 · Text
worldbank_project_documentsDataset Card for World Bank Project Documents Dataset Summary This is a dataset of documents related to World Bank development projects in the period 1947-2020. The dataset includes the documents used to propose or describe projects when they are launched, and those in the review. The documents are indexed by the World Bank project ID, which can be used to obtain features from multiple publicly av
Hugging Face Datasets2022 · dataset
GEMGEM is a benchmark environment for Natural Language Generation with a focus on its Evaluation, both through human annotations and automated Metrics. GEM aims to: - measure NLG progress across 13 datasets spanning many NLG tasks and languages. - provide an in-depth analysis of data and models presented via data statements and challenge sets. - develop standards for evaluation of generated text usin
Hugging Face Datasets2022 · Text
dartDART is a large and open-domain structured DAta Record to Text generation corpus with high-quality sentence annotations with each input being a set of entity-relation triples following a tree-structured ontology. It consists of 82191 examples across different domains with each input being a semantic RDF triple set derived from data records in tables and the tree ontology of table schema, annotated
Hugging Face Datasets2022 · Text
sportsett_basketballSportSett:Basketball dataset for Data-to-Text Generation contains NBA games stats aligned with their human written summaries.
Hugging Face Datasets2022 · Text
tottoToTTo is an open-domain English table-to-text dataset with over 120,000 training examples that proposes a controlled generation task: given a Wikipedia table and a set of highlighted table cells, produce a one-sentence description.
Hugging Face Datasets2022 · dataset
WikiBioThis dataset gathers 728,321 biographies from wikipedia. It aims at evaluating text generation algorithms. For each article, we provide the first paragraph and the infobox (both tokenized). For each article, we extracted the first paragraph (text), the infobox (structured data). Each infobox is encoded as a list of (field name, field value) pairs. We used Stanford CoreNLP (http://stanfordnlp.githu
Hugging Face Datasets2022 · dataset
ToTToToTTo is an open-domain English table-to-text dataset with over 120,000 training examples that proposes a controlled generation task: given a Wikipedia table and a set of highlighted table cells, produce a one-sentence description.
Hugging Face Datasets2022 · dataset
turku_hockey_data2textThe Turku Hockey Data2Text corpus was developed as a benchmark for evaluating template-free, machine learning methods on Finnish news generation in the area of ice hockey reporting. This dataset is a collection of 3,454 ice hockey games, each including game statistics and a news article describing the game. Each game includes manual alignment of events (such as goals or penalties) and sentences de
Hugging Face Datasets2022 · Text
RotoWire_English-GermanDataset for the WNGT 2019 DGT shared task on "Document-Level Generation and Translation”.
Hugging Face Datasets2022 · Text
e2e_nlgThe E2E dataset is designed for a limited-domain data-to-text task -- generation of restaurant descriptions/recommendations based on up to 8 different attributes (name, area, price range etc.).
Hugging Face Datasets2022 · Text
web_nlgWebNLG is a bi-lingual dataset (English, Russian) of parallel DBpedia triple sets and short texts that cover about 450 different DBpedia properties. The WebNLG data was originally created to promote the development of RDF verbalisers able to generate short text and to handle micro-planning (i.e., sentence segmentation and ordering, referring expression generation, aggregation); the goal of the tas
Hugging Face Datasets2022 · dataset
GREATThe dataset for the variable-misuse task, described in the ICLR 2020 paper 'Global Relational Models of Source Code' [https://openreview.net/forum?id=B1lnbRNtwr] This is the public version of the dataset used in that paper. The original, used to produce the graphs in the paper, could not be open-sourced due to licensing issues. See the public associated code repository [https://github.com/VHellend
Hugging Face Datasets2022 · Text
surface_realisation_st_2020Dataset Card for GEM/surface_realisation_st_2020 Link to Main Data Card You can find the main data card on the GEM Website. Dataset Summary This dataset was used as part of the multilingual surface realization shared task in which a model gets full or partial universal dependency structures and has to reconstruct the natural language. This dataset support 11 languages. You can load the dataset via
Hugging Face Datasets2022 · Text
viggoViGGO was designed for the task of data-to-text generation in chatbots (as opposed to task-oriented dialogue systems), with target responses being more conversational than information-seeking, yet constrained to the information presented in a meaning representation. The dataset, being relatively small and clean, can also serve for demonstrating transfer learning capabilities of neural models.
Hugging Face Datasets2022 · dataset
conversational_weatherThe Conversational Weather dataset is designed for generation of responses to weather queries based on a structured input data. The input allows specifying data attributes such as dates, times, locations, weather conditions, and errors, and also offers control over structure of response through discourse relations such as join, contrast, and justification.