Constarium
← Search

Data · dataset · 2013

Pan Matrix data

Listed in DataCite

The file pan_matrix.txt is a huge table (tab-separated columns) where each row corresponds to a genome and each column to a domain sequences family.

Description

The rows are named by the BIOID-code, see map_ecoli.txt to look up the strain names. The columns are named Cluster 1, Cluster 2,...etc.

The corresponding Pfam-A domain sequence is given in the file cluster_info.txt (see below). In cell (i,j) in this table you find the number of occurrences that domain sequence j has in genome number i.

Links

Topics

Stated by source
Biological sciences
Inferred from text
Tabular 65%
Provenance · 1 source records, 10 field assertions
SourceKeyLast seenRaw
DataCite10.6084/m9.figshare.10370712 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · DataCiteconnector:datacite@1.0.0/data/attributes/rightsList
byte_sizesource · DataCiteconnector:datacite@1.0.0
concepts[field].fos:biological-sciencessource · DataCiteconnector:datacite@1.0.0
concepts[modality].local:modality:tabularenrichment · DataCitekeyword-concept-rules@1.0.0title+description (65%)
created_datesource · DataCiteconnector:datacite@1.0.0
descriptionsource · DataCiteconnector:datacite@1.0.0/data/attributes/descriptions
licensesource · DataCiteconnector:datacite@1.0.0/data/attributes/rightsList
publication_datesource · DataCiteconnector:datacite@1.0.0/data/attributes/dates
titlesource · DataCiteconnector:datacite@1.0.0/data/attributes/titles/0/title
updated_datesource · DataCiteconnector:datacite@1.0.0