TWO ERAS · ONE RESEARCH TRAJECTORY

Image Retrieval.

My work in image retrieval spans two distinct technological eras: the early years of handcrafted visual descriptors and complete retrieval systems, and today’s shift toward deep networks, Vision Transformers, CLIP and universal image representations.

TODAY · 2021-PRESENT

From deep features to universal image embeddings.

These three articles document a continuous research path: from evaluating deep convolutional representations, to using Vision Transformers as image descriptors, and finally to learning universal CLIP embeddings across multiple visual domains. That progression led directly to the IRonCLIP solution for the Google Universal Image Embedding competition.

2023 · IEEE ACCESS

Universal Image Embedding.

Retaining and expanding knowledge with multi-domain fine-tuning.

A CLIP-based fine-tuning strategy that retains pre-existing knowledge while creating a domain-agnostic image encoder. Inspired by the team’s Kaggle result, the work develops the competition insight into a broader method that performs effectively on unseen retrieval and recognition domains.

S. Gkelios · A. Kastellos · Y. S. Boutalis · S. A. Chatzichristofis

Read the IEEE article

2021 · IEEE DCOSS

Investigating the Vision Transformer Model for Image Retrieval Tasks.

A plug-and-play global descriptor built from a pretrained Vision Transformer, with no task-specific training or fine-tuning. Across INRIA Holidays, UKBench, Paris6K and Oxford5K, it outperformed all evaluated handcrafted features and most CNN-based alternatives. This work established the transformer-based direction later advanced through IRonCLIP.

S. Gkelios · Y. Boutalis · S. A. Chatzichristofis · DCOSS 2021, pp. 367-373

Read the IEEE paper

2021 · EXPERT SYSTEMS WITH APPLICATIONS

Deep Convolutional Features for Image Retrieval.

A study of Inception-ResNet-v2, Xception, DenseNet201, MobileNet-v2 and EfficientNet as global and local image descriptors. It established the deep-feature baseline from which the later Vision Transformer and universal embedding research developed.

S. Gkelios · A. Sophokleous · S. Plakias · Y. Boutalis · S. A. Chatzichristofis

Read the journal article



RESEARCH OUTCOME · GOOGLE UNIVERSAL IMAGE EMBEDDING
IRonCLIP · Gold medal · 6th place
The CLIP and Vision Transformer solution developed by Socratis Gkelios, Anestis Kastellos and Savvas A. Chatzichristofis finished 6th among 1,022 teams, placing in the top 0.58% of the competition.

View leaderboard

THE EARLY YEARS · 2008-2019

Building the vocabulary of visual search.

Before learned representations became dominant, image retrieval depended on carefully designed visual features. This period focused on compact colour, texture, edge and spatial descriptors, then extended them into reusable libraries, desktop software and complete web retrieval systems.

THE EARLY YEARS · DESCRIPTORS AND LIBRARIES

The visual-description layer.

Methods and reusable components for representing colour, texture, spatial layout, edges and local image information.

COMPOSITE DESCRIPTORS

Compact Composite Descriptors

A family of compact descriptors combining colour and texture information for natural, medical and other specialised image collections. The page includes technical material, implementations and supporting resources.

Explore the descriptors

ISO/IEC 15938 · C#

MPEG-7 Visual Descriptors

Open-source C# implementations of the Scalable Color, Color Layout, Dominant Color and Edge Histogram descriptors, together with matching code and examples.

View implementations

LOCAL IMAGE DESCRIPTORS

SIMPLE Descriptors

A framework that applies MPEG-7 and related global descriptors to salient local patches generated through SURF, SIFT and random sampling strategies.

Explore SIMPLE

OPEN-SOURCE JAVA LIBRARY

LIRE

The Lucene Image Retrieval library creates indexes of visual features for content-based image retrieval. Its feature set includes MPEG-7 descriptors, autocorrelograms and additional research implementations.

Visit the LIRE project

INTEGRATED RESEARCH SUMMARY

Co.Vi.Wo.

Co.Vi.Wo., Color Visual Words Based on Non-Predefined Size Codebooks, extends the visual-words model by incorporating colour information directly into the construction of the visual vocabulary.

Instead of imposing a fixed codebook size in advance, the method allows the vocabulary to emerge from the visual and colour structure of the data. The resulting representation was designed to improve the descriptive power of image-retrieval systems while reducing dependence on manually selected codebook parameters.

S. A. Chatzichristofis · C. Iakovidou · Y. Boutalis · O. Marques · IEEE Transactions on Cybernetics, 2013

Read the research paper

THE EARLY YEARS · RESEARCH SOFTWARE

A complete retrieval laboratory.

C# · CBIR · RESEARCH ARCHIVE

img(Rummager)

An interactive platform combining established and original visual descriptors. It supports query-by-example search from XML indexes or folders, hybrid keyword and visual retrieval, feature extraction and standard retrieval-evaluation measures.

Explore img(Rummager)

THE EARLY YEARS · EXPERIMENTAL SYSTEMS

Prototypes preserved through their research record.

The original web demonstrations are presented here as historical systems. Their descriptions remain important as a record of the experimental questions, systems and evaluation tools developed around image and multimodal retrieval.

MULTIMODAL RETRIEVAL · ARCHIVE

MMRetrieval.net

An experimental multilingual and multimodal search engine. Modalities were indexed separately, while configurable fusion methods combined their ranked results.

View the research record

WEB CBIR · ARCHIVE

img(Anaktisi)

A web-based content-retrieval system built around compact colour and texture descriptors whose representations ranged from 23 to 74 bytes per image.

View the research paper