About

I'm a Research Scientist at Tavus, working on multimodal and generative human-centered AI systems.

I did my PhD in AI at the University of Amsterdam, working with Marcel Worring and Yuki Asano on multimodal foundation models, at the Informatics Institute. I was part of MultiX Amsterdam and AIMLab groups. During my PhD, I also spent time as a Research Scientist Intern in Meta, working on image generation and in-context learning.

I am a member of the ELLIS, co-organized WiML social and ICLR workshops, and served as a reviewer for leading conferences, including CVPR, ICCV, NeurIPS, ICLR, and ICML.

I obtained my master degree in Artificial Intelligence at the KU Leuven. Before that, I spent some time as a Software Engineer in Netcetera and I was an undergraduate student in Computer Science and Engineering at the FCSE at University โ€Ss. Cyril and Methodiusโ€ in Skopje.

News

  • Apr 2026I'm co-organizing the WiML social event at ICLR 2026 ๐ŸŽ‰
  • Dec 2025Our workshop Multimodal Intelligence: Next Token Prediction & Beyond will be part of ICLR 2026 ๐ŸŽ‰
  • Nov 2025๐ŸŽ‰ I successfully defended my PhD (thesis available here) ๐ŸŽ‰
  • Jul 2025The preprint for LATTEโ˜•๏ธ is available on arxiv.
  • Apr 2025I will be a TA for Foundation Models (FoMo) and Multimedia Analytics courses.
  • Mar 2025I gave a talk at Deepfakes & GenAI Workshop @ DEX-XL 2025.
  • Feb 2025TULIP๐ŸŒท is accepted to ICLR 2025 ๐ŸŽ‰
  • Oct 2024Context Diffusion is accepted to ECCV 2024 ๐ŸŽ‰
  • Sept 2024I got accepted to the ECCV 2024 Doctoral Consortium.
  • Apr 2024I will be a TA for the Foundation Models (FoMo) course at UvA.
  • Dec 2023The preprint of my internship work at Meta is available on arXiv.
  • June 2023I started a new position as a Research Scientist Intern at Meta AI in Menlo Park, California ๐Ÿ‡บ๐Ÿ‡ธ
  • Mar 2023I will be a TA for Deep Learning 2, Vision-Language learning module.
  • Jan 2023Our paper on multimodal few-shot learning is accepted to ICLR 2023 ๐ŸŽ‰
  • Nov 2022I taught a guest lecture on Attention & Transformers, as part of the Deep Learning 1 course.
  • Sept 2022Our paper is runner-up for the MEDIA Best Paper Award at MICCAI 2022

Research

My research centers around multimodal foundation models - with focus on designing efficient approaches for multimodal understanding and generative tasks. I'm interested in better understanding what large-scale models learn, and how to exploit that through in-context learning and prompting. Some of my work also includes automated linguistic interpretation of images, as well as its applications in the medical domain.

Publications (selected; full list on Google Scholar)

LATTE overview

LATTE: Latent Trajectory Embedding for Diffusion-Generated Image Detection

Ana Vasilcou*, Ivona Najdenkoska*, Zeno Geradts, Marcel Worring

Preprint (arXiv 2025) Paper Code

We present LATTE - Latent Trajectory Embedding - a novel approach for AI-generated image detection, which models the evolution of latent embeddings across several denoising timesteps.

ArtRAG overview

ArtRAG: Retrieval-Augmented Generation with Structured Context for Visual Art Understanding

Shuai Wang, Ivona Najdenkoska, Hongyi Zhu, Stevan Rudinac, Monika Kackovic, Nachoem Wijnberg, Marcel Worring

ACMMM 2025 Paper

We present ArtRAG - novel training-free framework that integrates structured knowledge into a RAG pipeline for multi-perspective artwork explanation.

SeCAt overview

Self-Supervised Open-Ended Classification with Small Visual Language Models

Mohammad M. Derakshani*, Ivona Najdenkoska*, Cees Snoek, Marcel Worring, Yuki M. Asano

ICLR ME-FoMo 2024 Paper

We present Self-Context Adaptation (SeCAt), a self-supervised approach that unlocks few-shot open-ended classification with small visual language models.

Meta-learned visual prefix model

Meta Learning To Bridge Vision and Language Models for Multimodal Few-Shot Learning

Ivona Najdenkoska, Xiantong Zhen, Marcel Worring

ICLR 2023 Paper Code

We propose a method for bridging large-scale vision and language models to perform multimodal few-shot learning. The model meta-learns visual prefixes from frozen visual backbone, which are used as prompts to a large language model.

Medical VQA model

Open-Ended Medical Visual Question Answering Through Prefix Tuning of Language Models

Tom van Sonsbeek*, Mohammad M. Derakshani*, Ivona Najdenkoska*, Cees Snoek, Marcel Worring

MICCAI 2023 Oral Paper Code

We introduce a novel method for open-ended VQA suited for small, domain-specific, medical datasets. We employ parameter-efficient strategies for efficient tuning of the LMs.

Meta-learning setting

Meta-Learning Makes a Better Multimodal Few-Shot Learner

Ivona Najdenkoska, Xiantong Zhen, Marcel Worring

NeurIPS 2022 Workshop on Meta-Learning Paper

We define a meta-learning approach for multimodal few-shot learning, to leverage its strong ability of accruing knowledge across tasks (predecessor of the ICLR 2023 work).

Uncertainty-aware report generation model

Uncertainty-aware Report Generation for Chest X-rays by Variational Topic Inference

Ivona Najdenkoska, Xiantong Zhen, Marcel Worring, Ling Shao

Medical Image Analysis 2022 Best Paper Honorable Mention Paper Code

We present a probabilistic latent variable model for chest X-Ray report generation. We extend the VTI model by providing a fully Transformer-based definition and explore the trade-off between an LSTM- or Transformer-based decoder for generation of medical text.

ECG captioning model

Learning to Automatically Generate Accurate ECG Captions

Mathieu G. G. Bartels, Ivona Najdenkoska, Rutger van de Leur, Arjan Sammani, Karim Taha, David M. Knigge, Pieter Doevendans, Marcel Worring, Rene van Es

MIDL 2022 Paper

We introduce a label-guided Transformer model, and show that it is possible to automatically generate relevant and readable ECG descriptions with a data-driven captioning model.

Variational Topic Inference model

Variational Topic Inference for Chest X-Ray Report Generation

Ivona Najdenkoska, Xiantong Zhen, Marcel Worring, Ling Shao

MICCAI 2021 Oral + Travel Award Paper Code

We propose Variational Topic Inference (VTI), a probabilistic latent variable model for automatic report generation. We introduce a set of topics as latent variables to guide sentence generation by aligning image and language modalities in the latent space.

Academic experience

University of Amsterdam ยท Teaching Assistant in Master of AI Feb 2021 โ€“ present

Work experience

  • June โ€“ Nov 2023Meta ยท Research Scientist Intern
  • Sept 2017 โ€“ Sept 2018Netcetera ยท Software Engineer
  • Apr 2017 โ€“ Jul 2017Netcetera ยท Software Engineering Intern
  • Jun 2016 โ€“ Sept 2016Haselt ยท Software Engineering Intern

Selected talks

  • Mar 2025Invited talk @ DEX-XL on Gen AI in Deepfake Detection, Noordwijkerhout
  • May 2024Invited talk @ NCCV 2024 about Context Diffusion paper, Den Bosch
  • Feb 2024Invited talk @ Core42 about Open-ended classification with small VLMs, Abu Dhabi
  • June 2023Invited talk @ Foundation Model lecture series by UvA x VU, Amsterdam
  • July 2022Invited talk @ Amsterdam Medical Data Science (AMDS) meetup, Amsterdam
  • July 2022Participant @ DeepLearn Summer School, organized by IRDTA, Gran Canaria
  • Aug 2021Participant @ Oxford Machine Learning Summer School, organized by AI for Global Goals, Oxford (virtual)