{
"version": "1.0",
"query": {
"text": null,
"operator": "and",
"filters": [
{
"field": "concept",
"op": "descendant_of",
"value": "ckan_group:rdm_inesctec_pt:cs-computer-science",
"evidence": null
}
],
"evidence_policy": "standard"
},
"sort": null,
"page": {
"size": 25
}
}INESC TEC Research Data Repository2026 · Image data
Taxonomy of SERP Elements Across Search EnginesThis dataset contains the dataset and supporting materials used in our research on Search Engine Results Pages (SERPs) structure, taxonomy, and element-level analysis across multiple search engines and languages. We used Python and Selenium Webdriver for web scraping the SERPs, ChatGPT-4o to translate the search queries. The dataset includes: - Multilingual search queries (English, Simplified Chin
INESC TEC Research Data Repository2026 · Text · unknown
LabadainLog-60: A Curated Tetun Query-Response Dataset for Conversational System Evaluation**1. Overview** LabadainLog-60 is a curated query dataset designed to evaluate and compare different conversational AI assistants for Tetun, derived from real user queries. It contains 60 queries selected from over 8,000 query log entries collected from two sources: - **Labadain Chat** (https://www.labadain.com): 2,101 entries collected from January 1–27, 2026 - **Labadain Old** (https://old.labad
INESC TEC Research Data Repository2025 · Text · unknown
CitiLink-Minutes: A Multilayer Annotated Dataset of Municipal Meeting Minutes (v1.0.0)CitiLink-Minutes (v1.0.0) is a multilayer annotated dataset of 120 municipal meeting minutes from Portuguese city councils (31k annotations: 20,375 entities, 11,163 relations). It includes three annotation layers: **Metadata** (participants, location, date, meeting type), **subjects of discussion**, and **voting information** (subjects, positions, and outcomes). Data is provided in JSON format, on
INESC TEC Research Data Repository2025 · dataset · unknown
ClaimPT: A Dataset for Claim Detection and Fact-CheckingClaimPT is a Portuguese dataset of manually annotated claims in news articles. It contains 1,308 articles provided by the LUSA news agency, each annotated by two trained linguists following a dedicated scheme for claim detection. Annotations include claim spans, claimers, topics, stance, and temporal information. The resource supports research in claim detection, fact-checking, and misinformation
INESC TEC Research Data Repository2025 · Table · CSV
Interface Element Frequencies in Search Engine Results Pages (SERPs) Across Query Intents, Search Engines and LanguagesThis dataset contains the data produced for the dissertation "User Interface Variations in Search Engine Results Pages Across Types of Search Queries and Search Engines". The project was conducted by student Adelaide Miranda Santos at FEUP, University of Porto, as part of the Masters in Informatics and Computing Engineering. The primary objective of this work is to study interface variations in se
INESC TEC Research Data Repository2025 · Table · CSV
LusoClin: Dataset of synthetic clinical notes in European Portuguese generated using an open-source large language model, along with prompting and evaluation dataLusoClin is a publicly available dataset of fully synthetic clinical notes in European Portuguese, generated using an open-source large language model and carefully curated prompts. The dataset simulates realistic clinical narratives while ensuring that no real patient data is included, enabling privacy-preserving research on clinical text retrieval. The primary purpose of LusoClin is to support t
INESC TEC Research Data Repository2024 · dataset
Artificial Intelligence and Infodemic: Video Dataset for Fact-Checked Health Communication and Synthetic MediaVideos created using prototypes and APIs for participatory research. The videos were used as technological probes presented to various stakeholders. This dataset was created in the context of Fact-Checking Chatbot Initiative. The proliferation of disinformation poses a significant challenge to societies. Within the field of journalism, fact-checking emerges as a critical tool to combat this issue.
INESC TEC Research Data Repository2024 · Image data
Semantic representation of the Registos de Baptismos da Paróquia de Aldoar (Porto, Portugal)This dataset comprises mappings of archival records from the National Archives of Portugal to the RiC-O (Records in Contexts Ontology) framework, namely the baptism registries of the Parish of Aldoar (Porto, Portugal) (PT/ADPRT/PRQ/PPRT01/001). The original EAD XML file represents the raw archival data, which was subsequently processed through a data extraction script. This extraction parsed the o
INESC TEC Research Data Repository2024 · Table · CSV
Wikipedia and Simple Wikipedia Lead Section Pairs for Nine CategoriesThe dataset (categorized_dataset folder) contains 9 files in .csv format, each a collection of 10,000 lead section pairs sourced from Wikipedia (https://www.wikipedia.org/) and Simple Wikipedia (https://simple.wikipedia.org/) for a given category. Included categories are Culture, Education, Employment, Entertainment, Health, Leisure, Objects, Science and Time. This dataset was created to understan
INESC TEC Research Data Repository2024 · Excel spreadsheet
Metadata and Analysis of Clinical Information Extraction Publications Using Large Language ModelsThis dataset contains all the data collected on all the papers analyzed in our publication, entitled "Harnessing Large Language Models for Clinical Information Extraction: A Systematic Literature Review", which is a systematic literature review of 85 publications on Clinical Information Extraction using LLMs, from 2019 to 2023. It includes metadata on every paper and explains our selection process
INESC TEC Research Data Repository2024 · Excel spreadsheet
Salary trends for public higher education teachers in Portugal (2004-2024)This dataset gathers data on the salaries of public higher education teachers in Portugal between 2004 and 2024.
INESC TEC Research Data Repository2024 · Excel spreadsheet
Tribunal do Santo Ofício in ArchOnto - Extension of archival records through Wikidata and DBpedia propertiesThis dataset comprises mappings of representations of archive records in ArchOnto, DBpedia, and Wikidata. These manual representations demonstrate how archive entities are depicted in linked data, individually within each model, and in a combination of all three. Through these representations, it becomes apparent which archive entities, as represented by ArchOnto, can be enhanced using properties
INESC TEC Research Data Repository2024 · Text
Labadain-30k+: A Monolingual Tetun Document-Level Audited DatasetLabadain-30k+ is a monolingual Tetun dataset containing 33,550 documents spanning from June 2001 to September 2023, excluding the years 2004 and 2005, for which no documents are available. Acquired through web crawling, the dataset is in text format and includes title, URL, source, category, publication date, and content. Each document is separated by two consecutive newlines.
INESC TEC Research Data Repository2024 · Table · CSV
Matrix profile analysis of Dansgaard-Oeschger events in palaeoclimate time seriesThis dataset includes all the datafiles and computational notebooks required to reproduce the work reported in the paper “Characterisation of Dansgaard-Oeschger events in palaeoclimate time series using the Matrix Profile”: ** Input datafiles** - time series (20-years resolution) of oxygen isotope ratios (δ18O) from NGRIP ice core on the GICC05 time scale (source: https://www.iceandclimate.nbi.ku.
INESC TEC Research Data Repository2023 · Excel spreadsheet
Automatic Quality Assessment of Wikipedia Articles - A Systematic Literature Review DatasetThis is the result dataset related to the article entitled "Automatic Quality Assessment of Wikipedia Articles - A Systematic Literature Review", which is a systematic literature review of 149 different papers related to the topic of automatic quality assessment of Wikipedia articles. The dataset includes metadata related to every assessed publication, providing full transparency over the entire p
INESC TEC Research Data Repository2023 · Excel spreadsheet
Research image management practices reported by scientific literature: Studies that explore the use and production of images in researchThe dataset includes a set of 109 articles that include relevant information about image management in the context of research. Selected articles are grouped according to Web of Science research domains.13 dimensions of analysis were created, in order to detail the characteristics of each of the works in relation to the effective management of images. Controlled vocabularies were also developed fo
INESC TEC Research Data Repository2023 · dataset
Representation of 25 records from the Portuguese National Archives in ArchontoThis dataset includes the manual representation of a sample of 25 archival records from the Portuguese National Archive in ArchOnto. The representation is based on the archive descriptions made available on DigitArq, the platform where archive records are made available, formulated considering the ISAD(G) standard. This representative sample was selected by a group of archive experts, considering
INESC TEC Research Data Repository2023 · dataset · unknown
Text2Story LusaThe Text2Story Lusa dataset contains 357 news articles published in European Portuguese by the Lusa news agency mostly between October 2020 and December 2020. The articles are in text format (.txt) and include: publication date, location, headline and content. Also included is a JSON file containing all news articles. This dataset was initially developed in the context of the project "Text2Story:
INESC TEC Research Data Repository2023 · Table · CSV
ArchOnto ontology representation of Portuguese archival description units (baptism records and passports).The content of the datasets include an excerpt of the ArchOnto ontology representation of the DigitArq baptisms records from Bragança District Archive and passports. These datasets, represented in ArchOnto, were obtained through an automatic tranformation from the CIDOC-CRM ontology representation. The initial records, represented in CIDOC-CRM, were the result of a migration process of records obt
INESC TEC Research Data Repository2023 · Excel spreadsheet
Analysis of the DigitArq archive record entities and their properties in Wikidata and in DBpediaThis dataset, which is the product of a master's dissertation (https://hdl.handle.net/10216/143034), analyzed a sample of 25 records from the Portuguese National Archives (https://digitarq.arquivos.pt/), chosen by archival specialists as representative of different fonds and description levels, to identify entities and properties and explore relationships with other non-archival resources. After s
INESC TEC Research Data Repository2023 · Excel spreadsheet
User evaluation of the DigitArq and DigitArq+ interfacesThis dataset is part of a dissertation entitled "Evaluation of the migration of archival records to linked data in the EPISA project" (https://hdl.handle.net/10216/143042 ). This dataset was generated in an experiment with two groups of users (common users and archivists) of both of the DigitArq platform interface and DigitArq+, a new interface developed in the EPISA project. A script with differe
INESC TEC Research Data Repository2023 · Archive
Consensual ArchOnto representation of 13 Portuguese Historical Archival Records based on their Digital RepresentationsThe dataset contains archival descriptions represented in the ArchOnto model (https://rdm.inesctec.pt/dataset/cs-2022-004) of 13 records from the 20th century with typewritten digital representations. We selected a sample of 13 records from a previously developed dataset (https://rdm.inesctec.pt/dataset/cs-2022-004) that extracted typewritten digital representations from the Arquivo Nacional da To
INESC TEC Research Data Repository2023 · Excel spreadsheet
Analysis of baptism, marriage and death registers belonging to the District Archives of Guarda, the District Archives of Porto and deferred passports from the District Archives of BragançaThe content of the datasets has description units related to baptisms, marriages and deaths from the District Archive of Guarda; baptisms, marriages and deaths from the District Archive of Porto; and, finally, deferred passports from the District Archive of Bragança. The analyzed records belonging to the District Archive of Guarda are present in a fonds called "Paróquia da Sé - Guarda 1860/1911-03
INESC TEC Research Data Repository2022 · Text
ISAD(G) Descriptions of Archival Records With Entity AnnotationThis dataset contains long text ISAD(G) fields from records from the Arquivo Nacional da Torre do Tombo annotated with entities. It was built to evaluate the effectiveness of data migration from ISAD(G) to ArchOnto in the EPISA project. Thus, a set of long text fields present in DigitArq was annotated, in order to evaluate the effectiveness of data migration from the old system to the new one. In
INESC TEC Research Data Repository2022 · dataset
Discourse Ontology for Incorporation/Transference EventsThe Discourse Ontology for Incorporation/Transference Events is an OWL ontology, a semantic interpretation ontology (SIO), that contains the concepts and relations used in natural language sentences describing these events. The ontology also contains a set of semantic web language rules (SWLR) to model dates. This dataset is part of the results obtained from the semantic migration process of Digit