Mohammad Albarham Mohammad Albarham

Structure from Motion Algorithm

2024

Structure from Motion (SfM) 3D Reconstruction Pipeline: Input Images, Feature Extraction, Image Matching, Camera Pose Estimation, 3D Triangulation, Ceres Solver Bundle Adjustment, and Final Reconstruction

Project information

MATLAB Computer Vision 3D Reconstruction SIFT RANSAC Bundle Adjustment Ceres Solver

About this project

This project is a MATLAB implementation of the Structure from Motion (SfM) algorithm. The SfM algorithm is a photogrammetric range imaging technique for estimating three-dimensional structures from two-dimensional image sequences that may be coupled with local motion signals. It is studied in the fields of computer vision and visual perception. The project is implemented using MATLAB and the code is available on both GitHub and MathWorks File Exchange. The project was a final project for Computer Vision Course at Chalmers University of Technology.

Pipeline Architecture & Methodology

Structure from Motion reconstructs a metric 3D scene from an unorganized collection of 2D photographs. As illustrated in the pipeline diagram above, the algorithm executes an end-to-end multi-view geometry workflow across seven core stages:

  • Input Images: Unordered collections of multi-view photographs capturing landmarks and architectural structures from varied distances, camera elevations, and viewing angles.
  • Feature Extraction: Detecting scale-invariant local interest points and computing 128-dimensional descriptors using the SIFT algorithm, providing robustness against scale differences and affine camera shifts.
  • Image Matching: Pairwise feature matching based on Euclidean descriptor distances with Lowe's ratio test to reject ambiguous and repetitive correspondence candidates.
  • Estimate Camera Poses: Estimating the Essential Matrix (E) from verified correspondences using the 8-point algorithm with RANSAC for outlier rejection, followed by Singular Value Decomposition (SVD) to recover relative camera rotations (R) and translation vectors (t).
  • Triangulate 3D Points: Calculating initial 3D scene point coordinates from verified 2D correspondences across multiple camera views via linear and nonlinear multi-view triangulation.
  • Bundle Adjustment: Formulating a large-scale nonlinear least-squares optimization problem solved via Ceres Solver, jointly refining all 3D scene coordinates and camera extrinsic/intrinsic parameters to minimize global reprojection errors.
  • Final Reconstruction: Visualizing the fully calibrated 3D point cloud alongside estimated camera positions and line-of-sight frustums.

Multi-Dataset Evaluation & Results

The algorithm was benchmarked across 9 diverse multi-view datasets featuring distinct geometric structures, scale variations, and camera trajectories.

Figure 1: 3D points of datasets 1 through 9. The arrangement follows a grid with 2 datasets per row.

Figure 1: 3D points of datasets 1 through 9. The arrangement follows a grid with 2 datasets per row, showing reconstructed point clouds and camera frustums.

Key observations across the experimental evaluations:

  • Planar & Architectural Facades (Datasets 1–3): Stable convergence along rectilinear camera baselines, accurately preserving geometric planar boundaries and structural wall profiles.
  • Orbital & Cylindrical Targets (Datasets 4–6): Consistent camera pose tracking along circular paths, demonstrating robust multi-view loop-closure stability around 360-degree structures.
  • Complex Volumetric Geometry (Datasets 7–9): High-density 3D point clouds capturing intricate architectural reliefs and structural depth with low residual reprojection error.

Key Features

Robust SIFT Matching

Scale-invariant feature detection paired with RANSAC-driven geometric filtering ensures dependable correspondence tracking even across wide baselines.

Epipolar Pose Estimation

Essential and Fundamental matrix estimation with SVD decomposition reliably recovers relative and absolute camera rotations and translations.

Ceres Solver Bundle Adjustment

Joint nonlinear least-squares optimization minimizes global reprojection error across all reconstructed 3D points and camera parameters.

9 Benchmark Datasets

Modular MATLAB architecture supporting automated execution and 3D visual comparison across 9 standard photogrammetric benchmark datasets.

Code Usage & API

The pipeline is executed in MATLAB via the run_sfm entry point:

% Run the Structure from Motion pipeline on a benchmark dataset
run_sfm(dataset_num, 'Visualize', true, 'LogLevel', 'verbose');

% Examples:
run_sfm(1);  % Reconstruct Dataset 1 (Leaning Tower of Pisa) with default settings
run_sfm(2, 'MaxIterations', 200, 'PlotCameras', true);  % Custom optimization settings

Technologies

  • MATLAB: Core pipeline implementation, matrix algebra, and 3D point cloud visualization
  • Ceres Solver: Large-scale nonlinear least-squares bundle adjustment
  • SIFT: Scale-Invariant Feature Transform for keypoint detection and description
  • RANSAC: Random Sample Consensus for robust outlier rejection in pose estimation
  • Multi-View Geometry: Epipolar constraints, Essential Matrix decomposition, and 3D triangulation