Epipolar geometry
View on Wikipedia
Two cameras take a picture of the same scene from different points of view. The epipolar geometry then describes the relation between the two resulting views.
Epipolar geometry is the geometry of stereo vision. When two cameras view a 3D scene from two distinct positions, there are a number of geometric relations between the 3D points and their projections onto the 2D images that lead to constraints between the image points. These relations are derived based on the assumption that the cameras can be approximated by the pinhole camera model.
Definitions
[edit]The figure below depicts two pinhole cameras looking at point X. In real cameras, the image plane is actually behind the focal center, and produces an image that is symmetric about the focal center of the lens. Here, however, the problem is simplified by placing a virtual image plane in front of the focal center i.e. optical center of each camera lens to produce an image not transformed by the symmetry. OL and OR represent the centers of symmetry of the two cameras lenses. X represents the point of interest in both cameras. Points xL and xR are the projections of point X onto the image planes.

Each camera captures a 2D image of the 3D world. This conversion from 3D to 2D is referred to as a perspective projection and is described by the pinhole camera model. It is common to model this projection operation by rays that emanate from the camera, passing through its focal center. Each emanating ray corresponds to a single point in the image.
Epipole or epipolar point
[edit]Since the optical centers of the cameras lenses are distinct, each center projects onto a distinct point into the other camera's image plane. These two image points, denoted by eL and eR, are called epipoles or epipolar points. Both epipoles eL and eR in their respective image planes and both optical centers OL and OR lie on a single 3D line.[1]
Epipolar line
[edit]The line OL–X is seen by the left camera as a point because it is directly in line with that camera's lens optical center. However, the right camera sees this line as a line in its image plane. That line (eR–xR) in the right camera is called an epipolar line. Symmetrically, the line OR–X is seen by the right camera as a point and is seen as epipolar line eL–xLby the left camera.
An epipolar line is a function of the position of point X in the 3D space, i.e. as X varies, a set of epipolar lines is generated in both images. Since the 3D line OL–X passes through the optical center of the lens OL, the corresponding epipolar line in the right image must pass through the epipole eR (and correspondingly for epipolar lines in the left image). All epipolar lines in one image contain the epipolar point of that image.[1] In fact, any line which contains the epipolar point is an epipolar line since it can be derived from some 3D point X.
Epipolar plane
[edit]As an alternative visualization, consider the points X, OL & OR that form a plane called the epipolar plane. The epipolar plane intersects each camera's image plane where it forms lines—the epipolar lines. The epipolar plane and all epipolar lines intersect the epipoles regardless of where X is located.
Epipolar constraint and triangulation
[edit]If the relative position of the two cameras is known, this leads to two important observations:
- Assume the projection point xL is known, and the epipolar line eR–xR is known and the point X projects into the right image, on a point xR which must lie on this particular epipolar line. This means that for each point observed in one image the same point must be observed in the other image on a known epipolar line. This provides an epipolar constraint: the projection of X on the right camera plane xR must be contained in the eR–xR epipolar line. All points X e.g. X1, X2, X3 on the OL–XL line will verify that constraint. It means that it is possible to test if two points correspond to the same 3D point. Epipolar constraints can also be described by the fundamental matrix,[1] or in the case of normalized image coordinates, the essential matrix[2] between the two cameras.
- If the points xL and xR are known, their projection lines are also known. If the two image points correspond to the same 3D point X the projection lines must intersect precisely at X. This means that X can be calculated from the coordinates of the two image points, a process called triangulation.[3]
Simplified cases
[edit]The epipolar geometry is simplified if the two camera image planes coincide. In this case, the epipolar lines also coincide (eL–XL = eR–XR). Furthermore, the epipolar lines are parallel to the line OL–OR between the centers of projection, and can in practice be aligned with the horizontal axes of the two images. This means that for each point in one image, its corresponding point in the other image can be found by looking only along a horizontal line. If the cameras cannot be positioned in this way, the image coordinates from the cameras may be transformed to emulate having a common image plane. This process is called image rectification.
Epipolar geometry of pushbroom sensor
[edit]In contrast to the conventional frame camera which uses a two-dimensional CCD, pushbroom camera adopts an array of one-dimensional CCDs to produce long continuous image strip which is called "image carpet". Epipolar geometry of this sensor is quite different from that of pinhole projection cameras. First, the epipolar line of pushbroom sensor is not straight, but hyperbola-like curve. Second, epipolar 'curve' pair does not exist.[4] However, in some special conditions, the epipolar geometry of the satellite images could be considered as a linear model.[5]
See also
[edit]Notes
[edit]- ^ a b c Hartley & Zisserman 2003, pp. 240-241
- ^ Hartley & Zisserman 2003, p. 257
- ^ Hartley & Zisserman 2003, p. 12
- ^ Jaehong Oh. "Novel Approach to Epipolar Resampling of HRSI and Satellite Stereo Imagery-based Georeferencing of Aerial Images" Archived 2012-03-31 at the Wayback Machine, 2011, accessed 2011-08-05.
- ^ Tatar, Nurollah; Arefi, Hossein (2019). "Stereo rectification of pushbroom satellite images by robustly estimating the fundamental matrix". International Journal of Remote Sensing. 40 (23): 8879–8898. Bibcode:2019IJRS...40.8879T. doi:10.1080/01431161.2019.1624862.
References
[edit]- Richard Hartley and Andrew Zisserman (2003). Multiple View Geometry in computer vision. Cambridge University Press. ISBN 0-521-54051-8.
Further reading
[edit]- Quang-Tuan Luong. "Learning Epipolar Geometry". Artificial Intelligence Center. SRI International. Archived from the original on 2021-06-28. Retrieved 2007-03-04.
- Robyn Owens. "Epipolar geometry". Retrieved 2007-03-04.
- Linda G. Shapiro and George C. Stockman (2001). Computer Vision. Prentice Hall. pp. 395–403. ISBN 0-13-030796-3.
- Vishvjit S. Nalwa (1993). A Guided Tour of Computer Vision. Addison Wesley. pp. 216–240. ISBN 0-201-54853-4.
- Roberto Cipolla and Peter Giblin (2000). Visual motion of curves and surfaces. Cambridge University Press, Cambridge. ISBN 0-521-63251-X.
Epipolar geometry
View on GrokipediaFundamental Concepts
Epipolar Plane
The epipolar plane is defined as the plane that contains the optical centers of two cameras, denoted as and , along with a point in the three-dimensional world, and is equivalently described as the plane spanned by the baseline—the line segment connecting and —and the point .[5] This plane plays a central role in epipolar geometry by establishing coplanarity among the camera centers, the world point, and the corresponding image points in each view.[5] Geometrically, the epipolar plane provides intuition for the constraints in stereo vision: the rays from to and from to both lie within this plane, and its intersection with the image planes of the two cameras produces a pair of corresponding lines—known as epipolar lines—that bound the possible locations of the projections of . This intersection limits the search for matching points between views to one dimension along these lines, rather than across the entire two-dimensional image, thereby simplifying correspondence problems in computer vision tasks.[5] The epipole in each image is the point where the baseline pierces the opposite image plane, serving as the intersection point for all such epipolar lines.[5] The concept of the epipolar plane originated in 19th-century studies of stereo vision and photogrammetry, with early formalization attributed to G. Hauck in 1883 that explored projective relations in paired images.[6] It was later rigorously integrated into modern computer vision through the foundational work of H.C. Longuet-Higgins in 1981, who developed algorithms for scene reconstruction that highlighted the plane's geometric implications for two-view correspondences.[7] Diagrams illustrating the epipolar plane typically depict the baseline as an axis around which a pencil of such planes rotates, with each plane slicing through the camera centers and a specific , showing the ray bundles from to each optical center confined within the plane and their projections forming intersecting lines on the image planes.[5] These visualizations emphasize how varying generates a family of epipolar planes, all sharing the baseline, to constrain multi-view projections.[5]Epipole
In epipolar geometry, the epipole refers to the projection of one camera's optical center onto the image plane of the other camera. Specifically, for two cameras with optical centers and , the epipole in the second image is the projection of , while in the first image is the projection of . This point arises from the epipolar plane formed by the baseline connecting the two optical centers and a point in the scene, where the epipole marks the intersection of the baseline with the respective image plane.[5] The epipole serves as a fixed point in each image that encapsulates the geometric relationship between the two views, independent of the scene structure. It represents the vanishing point of the baseline direction in the image and is the intersection of all epipolar lines corresponding to points along that baseline. In uncalibrated cameras, the epipole is the right null-vector of the fundamental matrix (satisfying ) or the left null-vector (satisfying ); in calibrated cameras, it similarly relates to the essential matrix through the calibration matrices and via .[5] A special case occurs when the cameras are in a parallel configuration involving pure translation parallel to the image planes, positioning the epipole at infinity and resulting in parallel epipolar lines across the images. Otherwise, the epipole occupies a finite position in the image plane, determined by the relative orientation and translation between the cameras.[5] Computationally, the epipole can be found geometrically by projecting the optical center of one camera using the projection matrix of the other; for instance, the epipole in the second image is given by , where is the projection matrix of the second camera and is the homogeneous coordinate vector of the first camera's center .[5]Epipolar Line
In epipolar geometry, the epipolar line in one image is defined as the intersection of the epipolar plane with that image's plane. This line represents the set of all possible projections onto the image plane of rays emanating from a 3D point through the center of the other camera.[5] A key property of epipolar lines is that all such lines in a given image pass through the epipole, which is the projection of the other camera's optical center. Corresponding points in the two images, say $ x_1 $ and $ x_2 $, lie on a pair of conjugate epipolar lines $ l_1 $ and $ l_2 $, ensuring that the rays from each camera center to these points are coplanar with the baseline connecting the centers.[5] Geometrically, the epipolar line plays a crucial role in stereo vision by constraining the search for matching points: instead of searching the entire 2D plane of the second image for a correspondent to a point in the first image, the search is reduced to a 1D line segment along the epipolar line. This simplification is fundamental to efficient stereo correspondence algorithms in computer vision.[5] For example, consider a point $ x_1 $ observed in the first image; the corresponding epipolar line $ l_2 $ in the second image is determined by the geometry of the two camera positions, forming a pencil of possible locations where the matching point $ x_2 $ must lie. In a typical diagram, this is illustrated as a line emanating from the epipole in the second image, highlighting the constraint imposed by the epipolar plane.[5]Mathematical Framework
Epipolar Constraint
The epipolar constraint is a fundamental algebraic relation in epipolar geometry that links corresponding points in two images captured by different cameras. For a 3D scene point projecting to homogeneous image points and in the first and second images, respectively, the constraint asserts that , where is the 3×3 fundamental matrix encoding the epipolar geometry between the views.[5] This constraint arises from the projective geometry of the pinhole camera model. Consider two cameras with projection matrices and , where and are intrinsic calibration matrices, is the relative rotation, and is the translation vector (baseline) between camera centers and . The projections are and , with denoting equality up to scale. The points , , and define an epipolar plane , which also contains the image points and . Geometrically, the optical rays from each camera center through and , together with the baseline, are coplanar. For the calibrated case where , the normalized coordinates satisfy , or equivalently with essential matrix . For uncalibrated cameras, incorporating intrinsics gives the fundamental matrix , leading to the bilinear form . This derivation holds under the assumption that the cameras are in general position, with non-coincident centers and no degenerate configurations.[5][2] Geometrically, the constraint enforces that the corresponding point must lie on the epipolar line in the second image, reducing the search space for matches from the entire plane to a one-dimensional line. This reflects the coplanarity of the optical rays from each camera center through and with the baseline. The derivation relies on the pinhole camera model, assuming ideal perspective projection without lens distortion or other aberrations, and requires the cameras to be uncalibrated or calibrated only through the intrinsics embedded in .[5] The fundamental matrix is a 3×3 homogeneous matrix of rank 2, possessing 7 degrees of freedom: it has 9 entries up to scale (8 DOF), minus one additional constraint from the rank deficiency . This singularity ensures that the epipole , the projection of the first camera center in the second image, satisfies , aligning with the geometric degeneracy at the epipole.[5]Fundamental Matrix
The fundamental matrix $ F $ is a 3×3 matrix that encodes the epipolar geometry between two uncalibrated cameras, satisfying the epipolar constraint $ \mathbf{x}'^\top F \mathbf{x} = 0 $ for corresponding points $ \mathbf{x} $ and $ \mathbf{x}' $ in the two images.[5] This matrix relates the projective structures of the two views without requiring knowledge of the camera intrinsics.[5] Key properties of $ F $ include its rank being exactly 2, which arises from the geometric constraint of the epipolar configuration.[5] The epipoles are the null vectors of $ F $ and $ F^\top $, satisfying $ F \mathbf{e} = 0 $ and $ F^\top \mathbf{e}' = 0 $, where $ \mathbf{e} $ and $ \mathbf{e}' $ are the epipoles in the respective images.[5] Epipolar lines are obtained as $ \mathbf{l}' = F \mathbf{x} $ in the second image for a point $ \mathbf{x} $ in the first, and symmetrically $ \mathbf{l} = F^\top \mathbf{x}' $ in the first image for a point $ \mathbf{x}' $ in the second.[5] Overall, $ F $ has 7 degrees of freedom due to its rank deficiency and scale ambiguity.[5] Estimation of $ F $ typically requires at least 8 point correspondences and employs the 8-point algorithm, which solves a linear system via least squares to minimize the algebraic error $ \sum (\mathbf{x}'^\top F \mathbf{x})^2 $.[8] The algorithm constructs a data matrix $ A $ from the correspondences, where each row encodes the outer product terms, and solves $ A \mathbf{f} = 0 $ for the vectorized $ F $ (denoted $ \mathbf{f} $) as the right singular vector corresponding to the smallest singular value of $ A $.[8] To enforce the rank-2 constraint post-estimation, singular value decomposition (SVD) is applied to set the smallest singular value to zero: if $ F = U \operatorname{diag}(\sigma_1, \sigma_2, \sigma_3) V^\top $, then the corrected $ F' = U \operatorname{diag}(\sigma_1, \sigma_2, 0) V^\top $.[8] For numerical stability, points are pre-normalized by translating to the origin (centroid) and scaling so the average distance from the origin is $ \sqrt{2} $, which significantly improves conditioning.[8] In the presence of outliers, the 8-point algorithm is often combined with RANSAC, which iteratively samples minimal subsets (8 points) to hypothesize $ F $, then counts inliers based on a thresholded epipolar error before refitting on the consensus set.[9][8] Decomposition of $ F $ allows recovery of camera matrices up to a projective transformation, with one canonical form being $ P = [I \mid 0] $ and $ P' = [\mathbf{e}'\times F \mid \mathbf{e}'] $, where $ \mathbf{e}' $ is the epipole and $ [\cdot]\times $ denotes the skew-symmetric matrix.[5] This reconstruction is ambiguous up to the 15 degrees of freedom in the projective group.[5] In degenerate cases, such as pure planar motion where all points lie on a plane, the fundamental matrix reduces in rank or structure, effectively relating to a homography between the views with 6 degrees of freedom.[5]Essential Matrix
The essential matrix, introduced by Christopher Longuet-Higgins in 1981, provides a fundamental constraint for calibrated stereo vision systems, enabling the recovery of relative camera pose from corresponding image points.[4] In the context of two calibrated cameras, it relates normalized image coordinates and (obtained by applying the inverse of the intrinsic matrix to pixel coordinates) through the equation .[10] This matrix encodes the epipolar geometry in Euclidean space and is expressed as , where is the 3×3 rotation matrix describing the orientation between the cameras, is the translation vector (up to scale), and denotes the skew-symmetric matrix formed from .[10] For uncalibrated cameras working with pixel coordinates and , the essential matrix relates to the fundamental matrix via , where and are the intrinsic calibration matrices of the respective cameras.[10] The essential matrix is a 3×3 matrix of rank 2, characterized by two equal non-zero singular values and one zero singular value, reflecting its geometric constraints.[10] It possesses 5 degrees of freedom, arising from the 3 degrees of freedom in the rotation matrix and the 2 degrees of freedom in the direction of the translation vector (with overall scale ambiguity).[2] To recover the camera pose, the essential matrix undergoes singular value decomposition (SVD) as , where with .[10] The decomposition proceeds by forming , yielding the translation direction as the last column of (up to sign) and rotation candidates via or , where is the matrix for a 90-degree rotation around the shared axis.[10] This process generates four possible relative pose configurations, and the correct one is selected through a chirality check, ensuring that the reconstructed 3D points lie in front of both cameras (positive depth).[10]Reconstruction and Applications
Triangulation
Triangulation in epipolar geometry recovers the 3D position of a point $ \mathbf{X} $ from its corresponding 2D projections $ \mathbf{x} $ and $ \mathbf{x}' $ in two calibrated images, by finding the intersection of the back-projected rays originating from the camera centers through these points along the epipolar lines. This process assumes known camera projection matrices $ P $ and $ P' $, and relies on the epipolar constraint to validate correspondences. The rays are typically parameterized in homogeneous coordinates as $ \mathbf{X} = \mathbf{C} + \lambda M^{-1} \tilde{\mathbf{x}} $ for the first camera, where $ \mathbf{C} $ is the camera center, $ M $ is the camera matrix, and $ \lambda $ is a scalar, with a similar form for the second view.[11][12] The linear method, known as the Direct Linear Transformation (DLT), solves the homogeneous system $ A \mathbf{X} = 0 $, where $ A $ is a $ 4 \times 4 $ matrix formed from the projection equations $ P \mathbf{X} = w \mathbf{x} $ and $ P' \mathbf{X} = w' \mathbf{x}' $ in homogeneous coordinates, with $ w $ and $ w' $ as scale factors. Specifically, the rows of $ A $ are derived from the cross-product forms $ \mathbf{x} \times (P \mathbf{X}) = 0 $ and $ \mathbf{x}' \times (P' \mathbf{X}) = 0 $, yielding two independent equations per view. The solution is obtained via singular value decomposition (SVD) of $ A $, taking $ \mathbf{X} $ as the right singular vector corresponding to the smallest singular value, ensuring $ | \mathbf{X} | = 1 $. This approach is projective-invariant and computationally efficient but minimizes algebraic error rather than geometric reprojection error. For refinement, non-linear least-squares optimization minimizes the reprojection error:function triangulateDLT(P, x, P_prime, x_prime):
# Form the 4x4 matrix A
A = zeros(4, 4)
A[0:2, :] = cross_product_matrix(x) * P
A[2:4, :] = cross_product_matrix(x_prime) * P_prime
# SVD: A = U * diag(s) * V^T
U, s, Vt = svd(A)
# X is the last column of V^T (smallest singular value)
X = Vt[-1, :] # [Homogeneous coordinates](/page/Homogeneous_coordinates)
# Normalize and check [chirality](/page/Chirality) (depth > 0)
if depth(P, X) > 0 and depth(P_prime, X) > 0:
return X / X[3] # Inhomogeneous
else:
return None # Invalid due to [chirality](/page/Chirality)
where cross_product_matrix computes the skew-symmetric matrix for the cross-product, and depth extracts the third component of the projected point.[11][12]