scale-invariant feature transform
Algorithm that detects and describes local image features in a way that is stable across changes in scale, rotation, and lighting. Used to identify keypoints for visual recognition tasks.
1 milestone
Bag of Words Applied to Computer Vision (Visual Vocabulary / Bag of Visual Words)
Josef Sivic and Andrew Zisserman at the University of Oxford applied the Bag of Words text-retrieval model to visual features in their 2003 ICCV paper 'Video Google', representing image regions as a vocabulary of visual words to enable efficient object retrieval from video.