Deep Learning-Driven Sparse Light Field Enhancement: A CNN-LSTM Framework for Novel View Synthesis and 3D Scene Reconstruction
Vivek Dwivedi, Gregor Rozinaj, Javlon Tursunov, Ivan Minárik, Marek Vanco, Radoslav Vargic
Sparse light field imaging often limits the quality of 3D scene reconstruction due to insufficient viewpoint coverage, resulting in incomplete or inaccurate reconstructions. This work introduces a hybrid CNN–LSTM-based framework to address this issue by generating novel camera poses and the corresponding synthesized novel views, effectively densifying the light field representation. The CNN extracts spatial features from the sparse input views, while the LSTM predicts temporal and positional dependencies, enabling smooth interpolation of novel poses and views. The proposed method integrates these synthesized views with the original sparse dataset to produce a comprehensive set of images. Our approach was evaluated on several datasets, including challenging datasets. The inference capability of our method was tested extensively, and it showed good generalization across diverse datasets. The effectiveness of the framework was evaluated not only with local light field fusion (LLFF) but also with NeRF and 3D Gaussian Splatting, which are considered state-of-the-art reconstruction methods. Overall, the enriched dataset generated by our method led to consistent improvements in 3D reconstruction quality, including higher depth estimation accuracy, reduced artifacts, and enhanced structural consistency. Most importantly, LSTM-based approaches have so far attracted limited attention in the context of generating novel views. While LSTMs have been widely applied in sequential data domains such as natural language processing, their use for image generation conditioned on camera poses remains largely unexplored, which underscores the novelty and significance of the proposed work. This approach provides a scalable and generalizable solution to the sparsity problem in light fields, advancing the capabilities of computational imaging, photorealistic rendering, and immersive 3D scene reconstruction. The results firmly establish the proposed method as a robust and versatile tool for improving reconstruction quality in sparse-view settings.