Constarium
← Search

Data · dataset · 2025

Anomaly Detection from ASRS Databases of Textual Reports

Listed in NASA Data Portal

Our primary goal is to automatically analyze textual reports from the Aviation Safety Reporting System (ASRS) database to detect/discover the anomaly categories reported by the pilots, and to assign each report to the appropriate category/categories.

Description

We have used two state-of-the-art models for text analysis: (i) mixture of von Mises Fisher (movMF) distributions, and (ii) latent Dirichlet allocation (LDA) on a subset of all ASRS reports.

The models achieve a reasonably high performance in discovering anomaly categories and clustering reports. Each category is represented by the most representative words with the highest probability in this category. In addition, since the inference algorithm for LDA was somewhat slow, we have developed a new fast LDA algorithm which is 5-10 times more efficient than the original one, therefore more applicable for the practical use.

Read the rest (1 more)

Further, we have developed a simple visualization tool based on non-linear manifold embedding (ISOMAP) to generate a 2-d visual representation of each report based on its content/topics, which gives a direct view of the structure of the whole dataset as well as the outliers.

Links

Topics

Inferred from text
Text 75%
Provenance · 1 source records, 7 field assertions
SourceKeyLast seenRaw
NASA Data Portaldbe51bf5-32b0-4a91-872e-8943b8ad742f10 d agoJSON v1
FieldAssertionExtractorEvidence
concepts[modality].local:modality:textenrichment · data nasa govkeyword-concept-rules@1.0.0title+description (75%)
created_datesource · data nasa govconnector:data_nasa_gov@1.0.0
descriptionsource · data nasa govconnector:data_nasa_gov@1.0.0/notes
license_textsource · data nasa govconnector:data_nasa_gov@1.0.0
publication_datesource · data nasa govconnector:data_nasa_gov@1.0.0
titlesource · data nasa govconnector:data_nasa_gov@1.0.0/title
updated_datesource · data nasa govconnector:data_nasa_gov@1.0.0