2009年3月3日

Scale & Affine Invariant Interest Point Detectors

Title:Scale & Affine Invariant Interest Point Detectors
Author: KRYSTIAN MIKOLAJCZYK AND CORDELIA SCHMID
Date of Publication: January 22, 2004

The paper describes two approaches for scale and affine invariant interest point detection. The scale invariant interest point detector, Harris-Laplace detector, combines the Harris detector with automatic scale selection. The algorithm involves a multi-scale point detection and an iterative selection of the scale and the location. The affine invariant interest point detector is initialized by the multi-scale Harris detector. Compute integration and differentiation scale to obtain shape matrix for each interest point. Finally, converge to a local structure in the iterative procedure.

What a pity is that there is a trade off between scale detection and affine detection because of different effects of the two kinds of detector.

Distinctive Image Featuresfrom Scale-Invariant Keypoints

Title:Distinctive Image Featuresfrom Scale-Invariant Keypoints
Author:David G. Lowe
Date of Publication: January 5, 2004

This paper describes an approach, Scale Invariant Feature Transform(SIFT), which transforms image data into scale-invariant coordinates relative to local features. There are several major stages of computation.

(1) The scale-space extrema detection stage searches all scales and image location and uses a cascade filtering approach to identify potential interest points that are invariant to scale and orientation. That is, compute the scale space of an image, a function produced from convolution of a variable-scale Gaussian with an input image, for scale feature description across all possible scales, and then use the function to compute the difference-of-Gaussian and find the maxima and minima of the DOG which can produce the most stable image features.

(2) The keypoint localization stage fits detailed model to determine location and scale at each candidate location and selects keypoints based on measures of their stability.

(3) The orientation assignment stage assigns one or more orientations to each keypoint location based on local image gradient directions; therefore, achieve invariance to image rotation.

(4) The keypoint descriptor stage measures the local image gradients at the selected scale in the region around each keypoint, which allows for significant levels of local shape distortion and change in illumination. In other words, a keypoint descriptor is created by computing the gradient magnitude and orientation at each image sample in a region. After weighted by a Gaussian window, the samples are accumulated into orientation histograms summarizing the contents over subregions. To improve the effects of shift and illumination change, trilinear interpolation, vector normalization and thresholding are taken into consideration.

In experiment on object recognition, an approximation algorithm, Best-Bin-First(BBF), is used for efficient nearest neighbor indexing to find minimum Euclidean distance in matching. Take Hough transform to cluster features. After solving affine parameters by least-square, accept a model if final probability is higher than a threshold.

Image Retrieval: Ideas, Influences, and Trends of the New Age

Title: Image Retrieval: Ideas, Influences, and Trends of the New Age
Authors: RITENDRA DATTA, DHIRAJ JOSHI, JIA LI, and JAMES Z. WANG
Year of Publication: 2008
Publisher: ACM

Content-based image retrieval is an technology helps to organize digital picture archives by their visual content. There are two gaps, sensory gap and semantic gap, which define and motivate most of the related problems. The sensory gap is a gap between real object and the descriptive information, and the semantic gap is about the lack of coincidence between the information from the visual data and the interpretation from a user. Therefore, how to solve the gaps and satisfy users is the goal of CBIR.

The first important thing is to clarify user-system interaction. A user perspective involves what the user wants and what is the form used in query, and a system perspective is about how to interact with user. Simply, human-center based system is required for different kinds of user intent with several query types, including keywords, free-text, image, graphics, and composite. For example, it is possible for a composite query method to provide a system involving gestures and speech for querying, or help user refine the queries by hints. If a system can collect manual tags for pictures, not only facilitating text-based querying, but also building reliable training datasets. Moreover, how to design a retrieval system on portable devices which have many constraints, such as limited size and color depth of display, is one of the issues of visualization.

The two core problems of CBIR are (1) how to define a mathematical description or a signature of an image, (2) how to decide the similarity between a pair of images.

For region-based visual signatures, the first step is image segmentation. With k-means clustering or normalized cut criteria method, segmentation helps image understanding and extracts several types of features. A feature capture a specific property of an image, either globally for the entire image with higher speed for computation, or locally for a small group of pixels with more specific identification of important visual characteristics. Color features are usually summarized into histogram. Texture features are used to capture granularity and repetitive patterns. Shape is a key to specify regions. Spatial modeling and matching are regarding to local image entities. Interest points that can deal with significant affine transformation and illumination changes are based on local invariants. When constructing signatures from features, histograms is easy but tend to be sparse in multidimensional space. A region-based signature allow representative vector to adapt images and the region of color and texture is likely corresponding to an object in an image.

There are three types of signatures, feature vector, summary of local feature vectors , and region-based signature. Each of them has different appropriate similarity measures. Using the geodesic distances for a single vector may be better. Summaries of local feature vectors such as codebook and probability density functions are generated by vector quantization and KL distance separately. The region-based signature can form a histogram, and calculate the similarity from the pair-wise distances between individual vectors. More matching methods improve the basic idea from region weights, speed, or segmentation.

Due to faster retrieval, clustering and classification is practical and useful. Classification is treated as a preprocessing step and improve accuracy but require prior training data. Clustering helps visualization and retrieval efficiency but may not representative enough or accurate for visualization. Besides, in order to capture user’s precise needs, relevance feedback system which does iterative feedback and refinement is designed. That is, relevance feedback let users give feedback after querying, and the system learns case-specific query semantics dynamically according to the feedback.

There are some offshoots of CBIR. First, automated annotation attempts at automated concept discovery; what is more is that deciding on an appropriate picture set for a given story. Second, ranking or similarity of images is usually sorted by size, color depth, or shape; however, aesthetics may be another higher-level basis which involves the feelings or emotions of people. Moreover, CBIR may concern with possible security attack or image copy protection.

Finally, evaluation benchmark of CBIR must some key points: coverage, unbiasedness, and user focus. Ideally, it should be subjective, context-specific, and community-based.