Accomplishments
- Multi-Floor Display
- Updated Experiences
- Mid-Year Review
Bonus
- Section Generation
- Camera Orientation Tools
- Local Language Automation
Halfway around the sun. That is how far we have travelled over the last 6 months. An incredible distance. And with this week, another incredible distance has been accomplished, this time in the world of spatial experiences. Each week has included accomplishments and research produced to detail and create processes and tools for efficiently delivering better spatial experiences.
Many improvements along the way contribute to this. Most notably, the User Experience, and the spatial reconstruction tool. While an incredible amount of improvements went into these, a meaningful summary can be found in my H1 video linked alongside this page.
This week brought H1 to the finish line, and delivered a meaningful end-to-end workflow alongside tools for delivering meaningful spatial experiences more efficiently, and affordably, than 6 months prior. To complete this chapter, I implemented a few final features. I improved the multi-floor display tool first drafted in week 2. I also defined section generation in the frontend, to pair with the responsive grid designed that same week. I created a new translator tool with a frontend that now simplifies the control over what gets translated and now automatically performs the translations using a local model, rather than calling an API, while still operating nicely with the structures defined in week 5. And to better tailor the visual experience, I created two new editors for setting the camera’s position when viewing the diorama, as well as each scene’s starting view.
Each new feature has been a step closer towards the original vision. Nothing is more telling than that than the fact that many of the features in the final week still cooperate with and improve upon the concepts defined in the early weeks.
Multi-Floor Display
One of the driving inspirations for this greater effort was determining how best to share large, multi-floor scenes. Many spaces have multiple levels, and the larger the spaces, the more they benefit from spatial awareness and the user being able to see where to go. Using a standard 3D model is often opaque and can be difficult to navigate. If you have a 3-floor home, and a walled-off room in the center of floor 2, it is very unlikely you can see or click it.
But what if the tour were smart enough to know when you want an unhindered view of the floor? When facing a floor on the horizontal plane, you often only see its walls, as the floor and ceiling are parallel to your vision. Here is a great opportunity to display other floors. Since you can’t see the ceiling or floor, levels above and below the selected one can fade in and become clickable to change the active floor. And when you angle up, the room’s floor comes into view and the walls fade away. These angles offer the opportunity to select locations, and require the other floors to disappear for focus and an uninterrupted click.
This result solves one of the bigger problems for a challenging tour, Supalai Place. This 3-story building with interiors and exteriors had many rooms, often ones which may be difficult to select if not for multi-floor display. With this feature, it is now much simpler to move quickly across the house. It reduces manually walking through, or reading the gallery, to just 2 clicks. One to the diorama, and one to the room.
Bonus: Section Generation
The responsive grid works great to simplify navigation and describe areas within a building alongside text descriptions. A challenge with it was manually defining a group and naming it for each scene. This is trivial for a dozen photos. It is much more cumbersome for a series of 100+ items, and much more prone to human error.
To solve this, I added a feature to the frontend where multiple images may be selected and set to a selection. These selections can then be exported alongside the other data and used when automatically writing the tour.xml file. This feature replaced manually placing keys throughout a file with a GUI-based definition including drag-selecting multiple cameras for a shared selection.
Bonus: Camera Placement Tools
A beautiful diorama is best seen from a good angle. If the camera places itself below or far away from the diorama, it may regularly cause users to have to move and zoom the camera to find a good location. This hinders the ability to navigate and appreciate a space. Also, when selecting a position, the camera’s angle may be unexpected. Staring at a blank wall or having a screen full of leaves may be momentarily jarring. And again, this causes the viewer to have to move and zoom in order to get an idea of what the scene is.
These were difficult to accept. Fixes do exist for them; manually calculating a good camera position and angle can be done, as well as manually copying coordinates for each photo’s camera to look at. This can add minutes or hours when creating an experience. With two new tools, they now take seconds per instance. Each diorama needs the camera to be placed in the desired spot, and at a click of a button the replacement data is available. The same goes for camera orientation per image. For each image, just adjust the camera, click a button, and its information is prepared.
Bonus: Local Language Automation
Multi-lingual support currently requires two steps: generating the translations and displaying them. The process was mostly formed in week 5. Displaying them has remained largely the same. The user selects a language and immediately all translated text is updated. However, under the hood, a lot has changed.
Previously, translations were done by manually typing a list of keys that were to be searched for and translated. Then, a paid API call was made to a machine learning tool in order to translate the files. And then new files were generated for each translation. This worked well to overcome some challenges in refreshing certain language text already loaded. Though, it came with its own consequences. This meant more redundant files to store and load. Loading a new series of files for every language change just didn’t feel necessary. And capable translators are available for download, so paying for the service seemed more wasteful than independent translation.
I found a local natural language ML model that is capable of good translations and set up a new UI to integrate with it. This UI gives the builder a filtered hierarchy of every available file, tag, and key for translation. The user can browse and select the ones important and begin the automation. And the process was further reworked to only need one new file per translation. The automated tool outputs a JSON file for each language with the translated files. A translation file will be loaded when chosen and used to replace the existing data.
Mid Year Review
Accomplishments
360 to 3D Reconstruction Tool
- Containerized Services
- Intuitive Frontend
- Masking
- Positioning
- Meshing
- Metric Depth Estimation
Automation
- Tripod Removal
- Hotspot Placement
- Tour File Building
- File Encryption
- Multi-Language Translation
UI Improvements
- Hide All Hotspots
- Information Hotspots
- 3D Cursor
- Deeplinking
- Media Player
Navigation
- Multi-Floor Fading
- 2D Custom Grid
- VR Custom Grid
- 3D Click and Go
Tools
- Multi-Crop Tool
- Thumbnail Updater
- Camera Orientation Tools
- Quantized Model for Mobile Inference
Capturing and sharing a moment is an art form as old as cave paintings and storytelling. The art form is often driven and pushed forward by the technological innovations of the time. Stone and dirt coated prehistoric caves. Cuneiform recorded knowledge that is still being reviewed today. Painting captured a vision and made it shareable. The camera captured a moment and made it copyable. Today, 360 cameras allow us to capture beyond a single perspective. Machine learning allows us to add depth to those images, and modern math allows us to calculate and pair them in 3D space. Together, these features represent the next step in sharing moments and recording knowledge. Unbound by the frame of a camera or the perspective of the photographer, spatial scenes let us teleport to and explore a moment.
Over the last 6 months, I have sought to identify a workflow for capturing, editing, and sharing spatial scenes. I reviewed existing processes, and I identified required features not yet present in a cohesive experience. I wanted developing scenes to be simple, fast, and cost-effective without compromising on quality or information density. To do this, I identified many tasks and set out a deadline. This required expanding upon existing tools while also creating my own.
Today, I share my achievements. Taking captured images and making them more useful has been simplified with image inpainting, which removes tripods automatically. A new system has been developed to then take these photos, estimate their depth, and combine them into a unified and navigable 3D model. To better orient oneself within an environment, I designed a user experience that pleasantly manages multi-floor scenes. And to better tailor the experience, I created a tool for translating tours into any language.
The user experience is the most important. The harder it is to interact with a tour, the less valuable it becomes to the viewer. A cave painting is simple; we just have to look at it. A photo is simple; we just have to look at it. And to look at the next one, we just have to flip the page. Spatial environments add a lot more opportunity, which can come with many more things for the viewer to learn. The current experience does its best to offer the viewer every opportunity in under 3 clicks.
With an improved user experience and newly developed tools to support it, I believe I have achieved each milestone and more during the first half of this year. I have since updated my existing spatial experiences to include these updates. I look forward to hearing your experiences. Thank you.


