TODAY · 2021-PRESENT
From deep features to universal image embeddings.
These three articles document a continuous research path: from evaluating deep convolutional representations, to using Vision Transformers as image descriptors, and finally to learning universal CLIP embeddings across multiple visual domains. That progression led directly to the IRonCLIP solution for the Google Universal Image Embedding competition.
2023 · IEEE ACCESS
Universal Image Embedding.
Retaining and expanding knowledge with multi-domain fine-tuning.
A CLIP-based fine-tuning strategy that retains pre-existing knowledge while creating a domain-agnostic image encoder. Inspired by the team’s Kaggle result, the work develops the competition insight into a broader method that performs effectively on unseen retrieval and recognition domains.
S. Gkelios · A. Kastellos · Y. S. Boutalis · S. A. Chatzichristofis
2021 · IEEE DCOSS
Investigating the Vision Transformer Model for Image Retrieval Tasks.
A plug-and-play global descriptor built from a pretrained Vision Transformer, with no task-specific training or fine-tuning. Across INRIA Holidays, UKBench, Paris6K and Oxford5K, it outperformed all evaluated handcrafted features and most CNN-based alternatives. This work established the transformer-based direction later advanced through IRonCLIP.
S. Gkelios · Y. Boutalis · S. A. Chatzichristofis · DCOSS 2021, pp. 367-373
2021 · EXPERT SYSTEMS WITH APPLICATIONS
Deep Convolutional Features for Image Retrieval.
A study of Inception-ResNet-v2, Xception, DenseNet201, MobileNet-v2 and EfficientNet as global and local image descriptors. It established the deep-feature baseline from which the later Vision Transformer and universal embedding research developed.
S. Gkelios · A. Sophokleous · S. Plakias · Y. Boutalis · S. A. Chatzichristofis
THE EARLY YEARS · 2008-2019
Building the vocabulary of visual search.
Before learned representations became dominant, image retrieval depended on carefully designed visual features. This period focused on compact colour, texture, edge and spatial descriptors, then extended them into reusable libraries, desktop software and complete web retrieval systems.
THE EARLY YEARS · DESCRIPTORS AND LIBRARIES
The visual-description layer.
Methods and reusable components for representing colour, texture, spatial layout, edges and local image information.
COMPOSITE DESCRIPTORS
Compact Composite Descriptors
A family of compact descriptors combining colour and texture information for natural, medical and other specialised image collections. The page includes technical material, implementations and supporting resources.
ISO/IEC 15938 · C#
MPEG-7 Visual Descriptors
Open-source C# implementations of the Scalable Color, Color Layout, Dominant Color and Edge Histogram descriptors, together with matching code and examples.
LOCAL IMAGE DESCRIPTORS
SIMPLE Descriptors
A framework that applies MPEG-7 and related global descriptors to salient local patches generated through SURF, SIFT and random sampling strategies.
OPEN-SOURCE JAVA LIBRARY
LIRE
The Lucene Image Retrieval library creates indexes of visual features for content-based image retrieval. Its feature set includes MPEG-7 descriptors, autocorrelograms and additional research implementations.
INTEGRATED RESEARCH SUMMARY
Co.Vi.Wo.
Co.Vi.Wo., Color Visual Words Based on Non-Predefined Size Codebooks, extends the visual-words model by incorporating colour information directly into the construction of the visual vocabulary.
Instead of imposing a fixed codebook size in advance, the method allows the vocabulary to emerge from the visual and colour structure of the data. The resulting representation was designed to improve the descriptive power of image-retrieval systems while reducing dependence on manually selected codebook parameters.
S. A. Chatzichristofis · C. Iakovidou · Y. Boutalis · O. Marques · IEEE Transactions on Cybernetics, 2013
THE EARLY YEARS · RESEARCH SOFTWARE
A complete retrieval laboratory.
C# · CBIR · RESEARCH ARCHIVE
img(Rummager)
An interactive platform combining established and original visual descriptors. It supports query-by-example search from XML indexes or folders, hybrid keyword and visual retrieval, feature extraction and standard retrieval-evaluation measures.
THE EARLY YEARS · EXPERIMENTAL SYSTEMS
Prototypes preserved through their research record.
The original web demonstrations are presented here as historical systems. Their descriptions remain important as a record of the experimental questions, systems and evaluation tools developed around image and multimodal retrieval.
MULTIMODAL RETRIEVAL · ARCHIVE
MMRetrieval.net
An experimental multilingual and multimodal search engine. Modalities were indexed separately, while configurable fusion methods combined their ranked results.
WEB CBIR · ARCHIVE
img(Anaktisi)
A web-based content-retrieval system built around compact colour and texture descriptors whose representations ranged from 23 to 74 bytes per image.