Data · dataset · 2026
TOWARD SPATIAL INTELLIGENCE AND PHYSICALLY REALISTIC WORLD MODELS
Listed in ZivaHub and Deakin Research Online and DMU Figshare — shown once because both records carry DOI 10.25394/pgs.34001037.v1
<p dir="ltr">Spatial intelligence requires models to infer and maintain coherent three-dimensional structure, preserve objects and relations across viewpoints, compose interactable environments, and respect geometric and physical constraints.
Description
Progress toward such world models is limited by two forms of data scarcity: a shortage of diverse, high-quality real-world spatial observations and the narrow distributions represented by curated interactive 3D scene datasets.
This dissertation studies how structured data and reusable priors can address both bottlenecks.</p><p dir="ltr">The first contribution, DL3DV-10K, establishes a high-quality corpus of 10,510 diverse real-world scenes and 51.2 million frames captured as multiview videos, together with a reproducible data-production protocol. Its benchmark reveals novel view synthesis failure modes missed by smaller collections; its scaling experiments show that broader real-world data improves generalizable 3D representation learning; and its adoption across 3D vision, generative video, and hybrid 3D--video systems demonstrates its value as shared infrastructure.</p><p dir="ltr">The second and third contributions study complementary routes to interactive 3D scene generation beyond curated distributions.
Read the rest (2 more)
Scenethesis grounds language planning and visual guidance with explicit 3D assets and geometric and physical constraints, producing editable indoor and outdoor scenes with more coherent relations, support, and collision behavior. I-Scene instead repurposes a pretrained 3D instance generator as a feed-forward scene-level spatial learner. Learning from non-semantic random compositions reduces its dependence on annotated semantic layouts and supports generalization to unseen object arrangements.</p><p dir="ltr">Together, these works show that progress toward spatially intelligent and physically realistic world models depends not only on model design but also on how spatial experience is collected, structured, grounded, and reused.
High-quality multiview data supplies transferable observations, while inference-time and learning-based approaches expand interactive 3D environment generation beyond bounded datasets. These findings provide foundations for persistent, interactive, and physically grounded world models.</p>
Links
Where it is published
- DOI doi.org/10.25394/pgs.34001037.v1 ↗
DOI / persistent id · from zivahub uct ac za
Catalogue records · 1
- OAI-PMH record api.figshare.com/v2/oai?verb=GetRecord&metadataPrefix=oai_dc&identifier=oai%3Af… ↗
metadata API · from zivahub uct ac za
Topics
- From keywords
- Autonomous agents and multiagent systems · Autonomous agents and multiagent systems · Autonomous agents and multiagent systems · Computer Science & AI · Computer Science & AI · Computer Science & AI · Computer vision · Computer vision · Computer vision · Earth & Environmental Science · Earth & Environmental Science · Earth & Environmental Science · Knowledge representation and reasoning · Knowledge representation and reasoning · Knowledge representation and reasoning · Modelling and simulation · Modelling and simulation · Modelling and simulation · Pattern recognition · Pattern recognition · Pattern recognition
- Inferred from text
- Video 75%
Provenance · 3 source records, 27 field assertions
| Source | Key | Last seen | Raw |
|---|---|---|---|
| ZivaHub | oai:figshare.com:article/34001037 | 5 d ago | JSON v1 |
| Deakin Research Online | oai:figshare.com:article/34001037 | 5 d ago | JSON v1 |
| DMU Figshare | oai:figshare.com:article/34001037 | 5 d ago | JSON v1 |
| Field | Assertion | Extractor | Evidence |
|---|---|---|---|
| access_level | source · zivahub uct ac za | connector:zivahub_uct_ac_za@1.0.0 | |
| concepts[field].anzsrc:field:460202 | mapping · figshare dmu ac uk | vocabulary-mapper@1.0.0 | keywords['Autonomous agents and multiagent systems'] |
| concepts[field].anzsrc:field:460202 | mapping · dro deakin edu au | vocabulary-mapper@1.0.0 | keywords['Autonomous agents and multiagent systems'] |
| concepts[field].anzsrc:field:460202 | mapping · zivahub uct ac za | vocabulary-mapper@1.0.0 | keywords['Autonomous agents and multiagent systems'] |
| concepts[field].anzsrc:field:460206 | mapping · figshare dmu ac uk | vocabulary-mapper@1.0.0 | keywords['Knowledge representation and reasoning'] |
| concepts[field].anzsrc:field:460206 | mapping · zivahub uct ac za | vocabulary-mapper@1.0.0 | keywords['Knowledge representation and reasoning'] |
| concepts[field].anzsrc:field:460206 | mapping · dro deakin edu au | vocabulary-mapper@1.0.0 | keywords['Knowledge representation and reasoning'] |
| concepts[field].anzsrc:field:460207 | mapping · dro deakin edu au | vocabulary-mapper@1.0.0 | keywords['Modelling and simulation'] |
| concepts[field].anzsrc:field:460207 | mapping · figshare dmu ac uk | vocabulary-mapper@1.0.0 | keywords['Modelling and simulation'] |
| concepts[field].anzsrc:field:460207 | mapping · zivahub uct ac za | vocabulary-mapper@1.0.0 | keywords['Modelling and simulation'] |
| concepts[field].anzsrc:field:460304 | mapping · figshare dmu ac uk | vocabulary-mapper@1.0.0 | keywords['Computer vision'] |
| concepts[field].anzsrc:field:460304 | mapping · zivahub uct ac za | vocabulary-mapper@1.0.0 | keywords['Computer vision'] |
| concepts[field].anzsrc:field:460304 | mapping · dro deakin edu au | vocabulary-mapper@1.0.0 | keywords['Computer vision'] |
| concepts[field].anzsrc:field:460308 | mapping · zivahub uct ac za | vocabulary-mapper@1.0.0 | keywords['Pattern recognition'] |
| concepts[field].anzsrc:field:460308 | mapping · figshare dmu ac uk | vocabulary-mapper@1.0.0 | keywords['Pattern recognition'] |
| concepts[field].anzsrc:field:460308 | mapping · dro deakin edu au | vocabulary-mapper@1.0.0 | keywords['Pattern recognition'] |
| concepts[field].local:field:computer-science-ai | mapping · dro deakin edu au | connector:dro_deakin_edu_au@1.0.0 | |
| concepts[field].local:field:computer-science-ai | mapping · figshare dmu ac uk | connector:figshare_dmu_ac_uk@1.0.0 | |
| concepts[field].local:field:computer-science-ai | mapping · zivahub uct ac za | connector:zivahub_uct_ac_za@1.0.0 | |
| concepts[field].local:field:earth-environmental | mapping · figshare dmu ac uk | connector:figshare_dmu_ac_uk@1.0.0 | |
| concepts[field].local:field:earth-environmental | mapping · zivahub uct ac za | connector:zivahub_uct_ac_za@1.0.0 | |
| concepts[field].local:field:earth-environmental | mapping · dro deakin edu au | connector:dro_deakin_edu_au@1.0.0 | |
| concepts[modality].local:modality:video | enrichment · zivahub uct ac za | keyword-concept-rules@1.0.0 | title+description (75%) |
| description | source · zivahub uct ac za | connector:zivahub_uct_ac_za@1.0.0 | /metadata/dc/description |
| license | source · zivahub uct ac za | connector:zivahub_uct_ac_za@1.0.0 | /metadata/dc/rights |
| publication_date | source · zivahub uct ac za | connector:zivahub_uct_ac_za@1.0.0 | |
| title | source · zivahub uct ac za | connector:zivahub_uct_ac_za@1.0.0 | /metadata/dc/title |