Applying neural reconstruction in simulation for automated driving - aiMotive

Applying neural reconstruction in simulation for automated driving

Validating automated driving software requires millions of test kilometers. This not only implies long system development cycles with continuously increasing complexity, but it also brings with it the problem that real-world testing is resource intensive, and safety issues might arise as well. A virtual validation suite like aiSim can alleviate these burdens of real-world testing.

Automated driving (AD) and Advanced Driver Assistance Systems (ADAS) rely on closed-loop validation to ensure safety and performance. However, achieving closed-loop evaluation requires a 3D environment that accurately represents real-world scenarios. While these 3D environments can be built manually by 3D artists, these solutions have limitations in scalability and addressing the Sim2Real domain gap.

Neural Rendering – Bridging the Gap

Neural rendering can mitigate this issue by leveraging deep learning techniques, it can realistically render the static (and dynamic) environments from novel viewpoints. Let's explore the pros and cons of this approach:

PROS:

CONS:

Challenges of Existing Generative Models

Current generative models are capable of creating highly realistic images and videos, but they fall short in several aspects, such as:

  1. 2D-Only Information: These models do not provide 3D information, only operate in a 2D image space.
  2. Projective Geometry Ignorance: They do not know projective geometry (see more HERE)
  3. Limited Sensor Modalities: These models can not be used for generating other sensor modalities (e.g., LiDAR).

In summary, the current generative models are NOT suitable for automotive-grade validation.

Our Hybrid Solution: Integrating Neural Reconstruction

To address these limitations, we at aiMotive developed a hybrid approach. The integration of cutting-edge neural reconstruction techniques within a well-established physically based rendering pipeline allows us to virtually insert dynamic objects at arbitrary locations, adjust environmental conditions, and render previously unseen camera viewpoints.

This way, we get the following features:

  1. Virtual Dynamic Content Insertion: – Add dynamic objects with realistic lighting and ambient occlusion. – Simulate environmental effects like rain, snow, and fog for more diverse simulated scenarios.

  2. Multi-Modality Renderings: – Generate accurate RGB images, depth-, and LiDAR intensity maps from arbitrary camera viewpoints (as seen below, with GT in the first row). – Future work is semantic segmentation masks and radar simulation.

  3. Camera Virtualization: – Simulate various virtual camera setups, including different camera alignments and models. – The figure below shows simulated front fisheye (left), front wide angle (middle), and front long-range (right) camera renders from a model trained without direct front cameras.