UC Berkeley Library Dataverse2026 · dataset · unknown
News on the WebThe NOW corpus (News on the Web) has been created by Mark Davies, and it contains 25.8 billion words of data from web-based newspapers and magazines from 2010 to the present time. More importantly, the corpus grows by about 240-250 million words of data each month (from about 400,000 new articles), or about three billion words each year. While other resources like Google Trends show you what peopl
UC Berkeley Library Dataverse2026 · dataset · unknown
Historical Business FilesConsists of annual snapshots of historic company data from Infogroup's U.S. Business Database (1997-2025). Includes company name, mailing address, SIC and NAICS codes, employee size, sales volume, latitude/longitude, and many more variables about each company.
UC Berkeley Library Dataverse2026 · dataset · unknown
MATERIAL Swahili-English Language PackMATERIAL Swahili-English Language Pack was developed by Appen for the IARPA (Intelligence Advanced Research Projects Activity) MATERIAL (Machine Translation for English Retrieval of Information in Any Language) program. It contains approximately 112 hours of Swahili conversational telephone speech, transcripts, English translations, annotations and queries. The MATERIAL program focused on underser
UC Berkeley Library Dataverse2026 · dataset · unknown
AnnoDIFP CTS Audio and TranscriptsAnnoDIFP (Annotated Data for the Investigation of Facets of Personality) CTS (Conversational Telephone Speech) Audio and Transcripts was developed by the Linguistic Data Consortium (LDC), the Florida Institute of Technology and the University of New Haven to support algorithm development for predicting personality traits. It contains 242.52 hours of English audio and transcripts from 1,179 telepho
UC Berkeley Library Dataverse2026 · dataset · unknown
2021 NIST Speaker Recognition Evaluation Development and Test Set2021 NIST Speaker Recognition Evaluation Test Set was developed by the Linguistic Data Consortium (LDC) and NIST (National Institute of Standards and Technology). It contains approximately 447 hours of Cantonese, Mandarin, and English conversational telephone speech (CTS), audio from video (AfV), and image data for development and test, along with answer keys, enrollment, trial files and documenta
UC Berkeley Library Dataverse2026 · dataset · unknown
LORELEI Russian Representative Language PackLORELEI Russian Representative Language Pack consists of Russian monolingual text, Russian-English parallel text, annotations, supplemental resources and related software tools developed by the Linguistic Data Consortium for the DARPA LORELEI program. The LORELEI (Low Resource Languages for Emergent Incidents) program was concerned with building human language technology for low resource languages
UC Berkeley Library Dataverse2026 · dataset · unknown
KAIROS Schema Learning Background Source DataKAIROS Schema Learning Background Source Data was developed by the Linguistic Data Consortium (LDC). It contains over 14,000 English and Spanish documents representing text, audio, video, image, and multimedia resources collected during the DARPA KAIROS program as supplemental background source data for the KAIROS Schema Learning Corpus (SLC). The complete set of SLC background source data was com
UC Berkeley Library Dataverse2026 · dataset · unknown
DEFT Chinese and English Light and Rich ERE Parallel AnnotationDEFT Chinese and English Light and Rich ERE Parallel Annotation was developed by the Linguistic Data Consortium (LDC) and consists of 179 Chinese discussion forum documents and their English translations annotated for entities, relations, and events (ERE). DARPA's Deep Exploration and Filtering of Text (DEFT) program aimed to address remaining capability gaps in state-of-the-art natural language p
UC Berkeley Library Dataverse2026 · dataset · unknown
Database of Word Level Statistics - MandarinDatabase of Word Level Statistics - Mandarin was developed by The Hong Kong Polytechnic University. It provides lexical characteristics of a descriptive and statistical nature for words and nonwords of Mandarin Chinese. It is designed for researchers particularly concerned with language processing of isolated words. Invariant characteristics include each item's lexicality, sampa, pinyin, IPA trans
UC Berkeley Library Dataverse2026 · Astronomical catalogue · unknown
Web of Science XML DataThe Web of Science XML Data includes metadata from over 12,500 journals spanning over 250 science, social science and humanities disciplines. Conference proceedings and book metadata are also available. Data are available back to 1900 and include over 63 million article records and 1 billion cited references to date. Some key data elements: ORCID identifiers are included in over 6.2 million record
UC Berkeley Library Dataverse2026 · dataset · unknown
CALLHOME Japanese Lexicon Second EditionCALLHOME Japanese Lexicon Second Edition was developed by the Linguistic Data Consortium (LDC) and contains 80,688 Japanese words with morphological, phonological and stress information. This second edition updates file formats, directory structure and documentation. The first edition is available as CALLHOME Japanese Lexicon (LDC96L17). The CALLHOME series consists of telephone conversations, tra
UC Berkeley Library Dataverse2026 · dataset · unknown
LORELEI Sinhala Incident Language PackLORELEI Sinhala Incident Language Pack was developed by the Linguistic Data Consortium (LDC) and consists of approximately 8.1 million words of Sinhala monolingual text, 70,000 words of English monolingual text, 6.4 million words of parallel Sinhala-English text, and 50,000 words of data annotated for Entity Discovery and Linking and Situation Frames. It contains all of the text data, annotations,
UC Berkeley Library Dataverse2026 · dataset · unknown
LORELEI Ilocano Incident Language PackLORELEI Ilocano Incident Language Pack was developed by the Linguistic Data Consortium (LDC) and consists of approximately 8.9 million words of Ilocano monolingual text, 3.3 million words of English monolingual text, 3.2 million words of parallel Ilocano-English text, and 3 million words of data annotated for Entity Discovery and Linking and Situation Frames. It contains all of the text data, anno
UC Berkeley Library Dataverse2026 · dataset · unknown
CALLHOME German Lexicon Second EditionCALLHOME German Lexicon Second Edition was developed by the Linguistic Data Consortium (LDC) and contains 318,809 German words with morphological, phonological, stress and frequency information. This second edition updates file formats, directory structure and documentation. The first edition is available as CALLHOME German Lexicon (LDC97L18). The CALLHOME series consists of telephone conversation
UC Berkeley Library Dataverse2026 · dataset · unknown
Air Traffic Control CompleteThe Air Traffic Control Corpus (ATC0) is comprised of recorded speech for use in supporting research and development activities in the area of robust speech recognition in domains similar to air traffic control (several speakers, noisy channels, relatively small vocabulary, constrained languaged, etc.) The audio data is composed of voice communication traffic between various controllers and pilots
UC Berkeley Library Dataverse2026 · dataset · unknown
MATERIAL Tagalog-English Language PackMATERIAL Tagalog-English Language Pack was developed by Appen for the IARPA (Intelligence Advanced Research Projects Activity) MATERIAL (Machine Translation for English Retrieval of Information in Any Language) program. It contains approximately 100 hours of Tagalog conversational telephone speech, transcripts, English translations, annotations and queries. The MATERIAL program focused on underser
UC Berkeley Library Dataverse2026 · dataset · unknown
2022 NIST Language Recognition Evaluation Test and Development Sets2022 NIST Language Recognition Evaluation Test and Development Sets was developed by the Linguistic Data Consortium (LDC) and the National Institute of Standards and Technology (NIST). This release contains the test and development data, metadata, answer keys, and documentation for the 2022 NIST Language Recognition Evaluation (LRE22). The source speech data is comprised of approximately 222 hours
UC Berkeley Library Dataverse2026 · dataset · unknown
CALLHOME Spanish Second EditionCALLHOME Spanish Second Edition was developed by the Linguistic Data Consortium (LDC) and contains approximately 38 hours of speech from 120 unscripted telephone conversations between native Spanish speakers. This publication is a re-release of the original CALLHOME Spanish collection, combining CALLHOME Spanish Speech (LDC96S35) and CALLHOME Spanish Transcripts (LDC96T17), with additional transcr
UC Berkeley Library Dataverse2026 · dataset · unknown
CALLHOME Japanese Second EditionCALLHOME Japanese Second Edition was developed by the Linguistic Data Consortium (LDC) and contains approximately 49 hours of speech from 120 unscripted telephone conversations between native Japanese speakers. This publication is a re-release of the original CALLHOME Japanese collection, combining CALLHOME Japanese Speech (LDC96S37) and CALLHOME Japanese Transcripts (LDC96T18), with additional tr
UC Berkeley Library Dataverse2026 · dataset · unknown
Supplemental Material [Videos] -- 2D Active BubblesSupporting videos for 2D active bubbles study.
UC Berkeley Library Dataverse2026 · dataset · unknown
Consumer Complaint DatabaseOnly complaints sent to companies for response are eligible to be published and are only published after the company responds, confirming a commercial relationship or after 15 days, whichever comes first. The database generally updates daily. We do not publish complaints referred to other regulators, such as complaints about depository institutions with less than $10 billion in assets. We publish
UC Berkeley Library Dataverse2025 · dataset · unknown
ACE 2005 Multilingual Training CorpusACE 2005 Multilingual Training Corpus was developed by the Linguistic Data Consortium (LDC) and contains approximately 1,800 files of mixed genre text in English, Arabic, and Chinese annotated for entities, relations, and events. This represents the complete set of training data in those languages for the 2005 Automatic Content Extraction (ACE) technology evaluation. The genres include newswire, b
UC Berkeley Library Dataverse2025 · dataset · unknown
Institute of Governmental Studies (IGS) Poll, 2025-10-20The Berkeley IGS Poll is a periodic survey of California public opinion on important matters of politics, public policy, and public issues. The poll, which is disseminated widely, seeks to provide a broad measure of contemporary public opinion, and to generate data for subsequent scholarly analysis. For IGS Poll press releases please visit: eScholarship IGS Poll Press Releases
UC Berkeley Library Dataverse2025 · dataset · unknown
Alphabetic listing of major war supply contractsAlphabetic listing of major war supply contracts published by the Civilian production administration, Industrial statistics division, c. 1946, from the Bancroft Library. Cumulative June 1940 through September 1945., Volume 1 A-C. Cumulative June 1940 through September 1945., Volume 2 D-J. Cumulative June 1940 through September 1945., Volume 3 K-Rex. Cumulative June 1940 through September 1945., Vo
UC Berkeley Library Dataverse2025 · dataset · unknown
CA Small Literary Press List, 2025A list of small presses publishing in California during the 2024 and/or 2025 calendar years. The data focuses on name of press, location as city and as street address, as well as URL, content levels (i.e., children or adult books), and some genre information. List includes founding date when listed on publisher's "About" page.