Tissue deformation
A physical point moves on a non-rigid surface with no stable shape.
VISION × LANGUAGE × SURGERY
VL-SurgPT pairs surgical point trajectories with textual descriptions such as pulled, reflected, or occluded—then uses that context to make tracking more robust.
01 · THE PROBLEM
Surgical scenes break the assumptions of ordinary tracking. Tissue bends, tools cover targets, the camera moves, wet surfaces reflect light, and cautery fills the view with smoke.
A physical point moves on a non-rigid surface with no stable shape.
A tool can hide the target while the tissue continues moving underneath.
Rapid viewpoint changes create blur and misleading global motion.
Highlights move independently and resemble visual landmarks.
Contrast disappears as boundaries become hazy or temporarily invisible.
(688, 502)
02 · THE DATASET
Data were collected from da Vinci Xi robotic procedures. ICG fluorescence anchors the ground truth; two clinicians then label intermediate frames at 1 fps with both coordinates and point status.
ANNOTATION WORKFLOW
ICG markers provide reliable endpoints. Clinicians use EVA to trace points and assign semantic labels through each clip.
DATASET VIEW
The label travels with the point, describing whether it is clear, pulled, reflected, obscured, or outside the field of view.
03 · THE METHOD
The method extends Track-On with an attribute prediction head and cross-modal attention. At inference time, it predicts its own semantic labels—no manual text is required.
SEE
Track-On extracts query, dense frame, and coarse matching features.
DESCRIBE
A status head estimates point type and current visual condition.
REFINE
Text–vision attention produces an offset that corrects the initial position.
04 · THE RESULTS
TG-SurgPT outperforms eight vision-only baselines across the reported benchmark metrics while retaining practical processing speed.
QUALITATIVE TRACKING
Tissue tracking across adverse scenarios. TG-SurgPT is blue, Track-On is red, and MFT is green.
05 · RESOURCES
VL-SurgPT opens a path from purely geometric tracking toward context-aware surgical perception. Use the links below to inspect the official paper and available data.