COMMUNITY

BOARD

News

Regular Paper Accepted at the Outstanding International Conference British Machine Vision Conference (BMVC) 2026

Author
College of Software Convergence
Date
2026-09-10
View
97

Files


Regular Paper Accepted at the Outstanding International Conference British Machine Vision Conference (BMVC) 2026



Master's student Lee Ji-hoon (Advisor: Professor Park Un-sang)

The paper “PASS-3D: Pose-Anchored Sphere-Ray Shuttle 3D Reconstruction from Monocular Broadcast Badminton”, written by Lee Ji-hoon (master's, first author) and Professor Park Un-sang (corresponding author) of the Computer Vision and Image Processing Laboratory (CVIP), has been accepted for publication at the outstanding international conference British Machine Vision Conference (BMVC) 2026 (acceptance rate: 406/1448 = 28.0%).

Sport has lately seen active efforts to turn match footage into three-dimensional data for tactical analysis, training and broadcast content. In badminton, however, three-dimensional trajectory data for the shuttlecock remains scarce. Obtaining it requires a multi-camera measurement system, which is expensive and can only be operated in specific venues. Broadcast footage, by contrast, has accumulated at a far greater scale, covering many more players, venues and rallies, but it remains largely two-dimensional, so it yields no information about the height of the shuttlecock, its clearance over the net, the height of the impact point, its position along the depth of the court or its three-dimensional speed.

Reconstructing the three-dimensional trajectory of a shuttlecock from a single broadcast camera is not easy. In 1080p footage the shuttlecock spans no more than 5 to 12 pixels, which general-purpose depth estimation models struggle with, and a smash reaches 350 km/h, a regime in which air resistance rather than gravity dominates the motion. Above all, unlike table tennis or basketball, badminton has no floor-contact event between strokes, so there is no reference point from which distance can be fixed. MonoTrack, the existing monocular study, uses the player's foot position together with an assumed height of 2 m as its reference point and validated three-dimensional accuracy only on synthetic data, while other studies have relied on manual work in which a person marks the shuttlecock and its ground projection frame by frame.

Starting from that problem, the team proposes PASS-3D, a monocular three-dimensional reconstruction framework that uses the player's body as a metric-scale reference. Noting that at the moment of impact the shuttlecock must lie within the reach of the racket, the framework determines the three-dimensional position of the impact point by intersecting the shuttlecock ray back-projected from the image with a sphere of racket-reach radius centred on the player's wrist. Of the two candidates where the ray meets the sphere, the true impact point is chosen by consistency with the direction of the forearm; in geometrically unfavourable situations where the ray is nearly tangent to the sphere and the two candidates cannot be clearly separated, a low confidence is assigned so that the candidate's influence on the subsequent optimisation is adjusted automatically.

The impact points obtained this way serve as constraints on the start and end of each flight segment, and the interval between them is connected by a flight equation that includes air resistance. A per-segment drag coefficient and a sub-frame observation timing error are estimated jointly, which reduces the problem of motion blur and detection error in fast strokes being transferred into depth error. Because the accuracy of the reference points governs overall performance, the depth and height drift of the human pose estimation model during jumps was corrected using the anatomical ankle height and the assumption of ballistic motion. As a result, across 73 rallies the two-dimensional pose reprojection error fell from an average of 6.9 pixels to 0.74 pixels.

The team also built a two-view benchmark for evaluation. A main camera at a broadcast-like viewpoint and an auxiliary side camera record simultaneously, and triangulation produces real three-dimensional ground truth; the set covers 73 rallies and 1,196 strokes over roughly 30 minutes of footage. The reprojection error for the moving shuttlecock is at the pixel level. To the best of the team's knowledge, this is the first study of monocular badminton reconstruction to obtain real three-dimensional ground truth automatically, rather than relying on synthetic data or manual annotation, and to evaluate against it.

In the experiments PASS-3D recorded average errors of 3.5 cm laterally, 7.4 cm in height and 24.5 cm in depth against the two-view ground truth. On match footage from the public TrackNetV2 dataset the reprojection error was about 2.4 pixels, an order of magnitude below the end-to-end error of the existing MonoTrack, and even when both methods were given the same ground-truth detection and stroke information the error was three to four times lower. In addition, replacing the wrist-based reference point with the foot-based reference point of the existing approach increased the three-dimensional error by 75%, and considering only gravity while excluding air resistance increased the error by a factor of 4.5, confirming that these two elements are central to the performance.

This research shows that broadcast footage that has already accumulated can be turned into metric-scale three-dimensional data without multi-camera measurement equipment. The information a person has to supply is limited to the impact frames, the stroke type and the jump segments, and frame-by-frame annotation of the shuttlecock is not required, which points to the possibility of scaling up to large broadcast archives. The reconstructed trajectories provide measures that cannot be obtained from two-dimensional trajectories or pose information alone, such as the height of the shuttlecock, its clearance over the net and court positioning.

Lee Ji-hoon, the master's student who is first author of the paper, said: “The significance of this research is that it fills in the depth information that monocular footage lacks with the geometry and physics specific to the sport, rather than with a general-purpose model. Because badminton has no reference event like the bounce in table tennis, we chose an approach that takes the player's wrist and the reach of the racket at the moment of impact as the reference. We also wanted to build an evaluation environment in which the results could be verified against real three-dimensional ground truth rather than synthetic data. Going forward we intend to extend the validation to a wider range of venues and players, and to continue the research towards using the reconstructed trajectories in real tactical analysis. I am grateful to Professor Park Un-sang for his guidance.”

The British Machine Vision Conference (BMVC), hosted by the British Machine Vision Association (BMVA), is a leading international conference in machine vision, image processing and pattern recognition. It is classified as an outstanding conference in the AI and computer vision field on the Korean Institute of Information Scientists and Engineers' list of outstanding conferences, and outstanding papers are invited to submit to a special issue of the International Journal of Computer Vision (IJCV). The 37th BMVC 2026 will be held in Lancaster, United Kingdom, from 23 to 26 November this year.

References: