Constarium
← Search

Data · dataset · 2026

TOWARD SPATIAL INTELLIGENCE AND PHYSICALLY REALISTIC WORLD MODELS

Listed in ZivaHub and Deakin Research Online and DMU Figshare — shown once because both records carry DOI 10.25394/pgs.34001037.v1

<p dir="ltr">Spatial intelligence requires models to infer and maintain coherent three-dimensional structure, preserve objects and relations across viewpoints, compose interactable environments, and respect geometric and physical constraints.

Description

Progress toward such world models is limited by two forms of data scarcity: a shortage of diverse, high-quality real-world spatial observations and the narrow distributions represented by curated interactive 3D scene datasets.

This dissertation studies how structured data and reusable priors can address both bottlenecks.</p><p dir="ltr">The first contribution, DL3DV-10K, establishes a high-quality corpus of 10,510 diverse real-world scenes and 51.2 million frames captured as multiview videos, together with a reproducible data-production protocol. Its benchmark reveals novel view synthesis failure modes missed by smaller collections; its scaling experiments show that broader real-world data improves generalizable 3D representation learning; and its adoption across 3D vision, generative video, and hybrid 3D--video systems demonstrates its value as shared infrastructure.</p><p dir="ltr">The second and third contributions study complementary routes to interactive 3D scene generation beyond curated distributions.

Read the rest (2 more)

Scenethesis grounds language planning and visual guidance with explicit 3D assets and geometric and physical constraints, producing editable indoor and outdoor scenes with more coherent relations, support, and collision behavior. I-Scene instead repurposes a pretrained 3D instance generator as a feed-forward scene-level spatial learner. Learning from non-semantic random compositions reduces its dependence on annotated semantic layouts and supports generalization to unseen object arrangements.</p><p dir="ltr">Together, these works show that progress toward spatially intelligent and physically realistic world models depends not only on model design but also on how spatial experience is collected, structured, grounded, and reused.

High-quality multiview data supplies transferable observations, while inference-time and learning-based approaches expand interactive 3D environment generation beyond bounded datasets. These findings provide foundations for persistent, interactive, and physically grounded world models.</p>

Links

Where it is published

Catalogue records · 1

Topics

Inferred from text
Video 75%
Provenance · 3 source records, 27 field assertions
SourceKeyLast seenRaw
ZivaHuboai:figshare.com:article/340010375 d agoJSON v1
Deakin Research Onlineoai:figshare.com:article/340010375 d agoJSON v1
DMU Figshareoai:figshare.com:article/340010375 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · zivahub uct ac zaconnector:zivahub_uct_ac_za@1.0.0
concepts[field].anzsrc:field:460202mapping · figshare dmu ac ukvocabulary-mapper@1.0.0keywords['Autonomous agents and multiagent systems']
concepts[field].anzsrc:field:460202mapping · dro deakin edu auvocabulary-mapper@1.0.0keywords['Autonomous agents and multiagent systems']
concepts[field].anzsrc:field:460202mapping · zivahub uct ac zavocabulary-mapper@1.0.0keywords['Autonomous agents and multiagent systems']
concepts[field].anzsrc:field:460206mapping · figshare dmu ac ukvocabulary-mapper@1.0.0keywords['Knowledge representation and reasoning']
concepts[field].anzsrc:field:460206mapping · zivahub uct ac zavocabulary-mapper@1.0.0keywords['Knowledge representation and reasoning']
concepts[field].anzsrc:field:460206mapping · dro deakin edu auvocabulary-mapper@1.0.0keywords['Knowledge representation and reasoning']
concepts[field].anzsrc:field:460207mapping · dro deakin edu auvocabulary-mapper@1.0.0keywords['Modelling and simulation']
concepts[field].anzsrc:field:460207mapping · figshare dmu ac ukvocabulary-mapper@1.0.0keywords['Modelling and simulation']
concepts[field].anzsrc:field:460207mapping · zivahub uct ac zavocabulary-mapper@1.0.0keywords['Modelling and simulation']
concepts[field].anzsrc:field:460304mapping · figshare dmu ac ukvocabulary-mapper@1.0.0keywords['Computer vision']
concepts[field].anzsrc:field:460304mapping · zivahub uct ac zavocabulary-mapper@1.0.0keywords['Computer vision']
concepts[field].anzsrc:field:460304mapping · dro deakin edu auvocabulary-mapper@1.0.0keywords['Computer vision']
concepts[field].anzsrc:field:460308mapping · zivahub uct ac zavocabulary-mapper@1.0.0keywords['Pattern recognition']
concepts[field].anzsrc:field:460308mapping · figshare dmu ac ukvocabulary-mapper@1.0.0keywords['Pattern recognition']
concepts[field].anzsrc:field:460308mapping · dro deakin edu auvocabulary-mapper@1.0.0keywords['Pattern recognition']
concepts[field].local:field:computer-science-aimapping · dro deakin edu auconnector:dro_deakin_edu_au@1.0.0
concepts[field].local:field:computer-science-aimapping · figshare dmu ac ukconnector:figshare_dmu_ac_uk@1.0.0
concepts[field].local:field:computer-science-aimapping · zivahub uct ac zaconnector:zivahub_uct_ac_za@1.0.0
concepts[field].local:field:earth-environmentalmapping · figshare dmu ac ukconnector:figshare_dmu_ac_uk@1.0.0
concepts[field].local:field:earth-environmentalmapping · zivahub uct ac zaconnector:zivahub_uct_ac_za@1.0.0
concepts[field].local:field:earth-environmentalmapping · dro deakin edu auconnector:dro_deakin_edu_au@1.0.0
concepts[modality].local:modality:videoenrichment · zivahub uct ac zakeyword-concept-rules@1.0.0title+description (75%)
descriptionsource · zivahub uct ac zaconnector:zivahub_uct_ac_za@1.0.0/metadata/dc/description
licensesource · zivahub uct ac zaconnector:zivahub_uct_ac_za@1.0.0/metadata/dc/rights
publication_datesource · zivahub uct ac zaconnector:zivahub_uct_ac_za@1.0.0
titlesource · zivahub uct ac zaconnector:zivahub_uct_ac_za@1.0.0/metadata/dc/title