<?xml version="1.0" encoding="utf-8"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
	<channel>
		<atom:link href="https://www.marcogarosi.it/blog/x5feed.php" rel="self" type="application/rss+xml" />
		<title><![CDATA[]]></title>
		<link>https://www.marcogarosi.it/blog/</link>
		<description><![CDATA[]]></description>
		<language>EN</language>
		<lastBuildDate>Thu, 26 Feb 2026 21:26:00 +0100</lastBuildDate>
		<generator>Incomedia WebSite X5 Pro</generator>
		<item>
			<title><![CDATA[CIRCLE available on arXiv]]></title>
			<author><![CDATA[Marco Garosi]]></author>
			<category domain="https://www.marcogarosi.it/blog/index.php?category=Research"><![CDATA[Research]]></category>
			<category>imblog</category>
			<description><![CDATA[<div id="imBlogPost_000000019"><div class="imHeading1">Read the full paper: our preprint is now on arXiv</div><div>Just a quick update: the full text of our CVPR Findings 2026 paper, <a href="https://www.marcogarosi.it/circle.html" class="imCssLink" onclick="return x5engine.utils.location('https://www.marcogarosi.it/circle.html', null, false)">"Large Multimodal Models as General In-Context Classifiers"</a>, is now officially available as a <a href="https://arxiv.org/abs/2602.23229" target="_blank" class="imCssLink">preprint on arXiv</a>.</div><div>If you want to dive straight into the technical details behind our benchmarking, the open-world challenges of large multimodal models (LMMs), and the exact mechanics of our <b>training-free <i>CIRCLE</i> method</b>, the preprint is already available!</div></div>]]></description>
			<pubDate>Thu, 26 Feb 2026 20:26:00 GMT</pubDate>
			<enclosure url="https://www.marcogarosi.it/blog/files/circle-on-arxiv_thumb.webp" length="8804655" type="image/webp" />
			<link>https://www.marcogarosi.it/blog/?circle-available-on-arxiv</link>
			<guid isPermaLink="false">https://www.marcogarosi.it/blog/rss/000000019</guid>
		</item>
		<item>
			<title><![CDATA[Project page and code for CIRCLE are available]]></title>
			<author><![CDATA[Marco Garosi]]></author>
			<category domain="https://www.marcogarosi.it/blog/index.php?category=Research"><![CDATA[Research]]></category>
			<category>imblog</category>
			<description><![CDATA[<div id="imBlogPost_000000018"><div class="imHeading1">Code and project page now live: exploring LMMs for classification with CIRCLE</div><div>Following up on the recent <a href="https://www.marcogarosi.it/blog/?circle-has-been-accepted-to-cvpr-findings" class="imCssLink">acceptance of our paper to <b data-path-to-node="3" data-index-in-node="54">CVPR Findings 2026</b></a>, I am excited to share that the official project page and the full codebase for our research are now publicly available!</div><div><br></div><div>If you read my previous post, you know we have been investigating how large multimodal models (LMMs) can challenge contrastive vision-language models (VLMs) like CLIP in classification tasks. By leveraging in-context learning, we found that LMMs can serve as powerful, unified classifiers - especially in open-world scenarios.</div><div>To make our findings as accessible and reproducible as possible, we have open-sourced everything.</div><hr data-path-to-node="6"><div class="imHeading2">Explore the project page</div><div>We have set up a <a href="https://circle-lmm.github.io/" target="_blank" class="imCssLink">dedicated project page</a>. Take a look at it - it has some cool animations that help you understand the method!</div><div class="imHeading2">Get hands-on with the code on GitHub</div><div>We also released the code on GitHub, at <a href="https://github.com/marco-garosi/CIRCLE" target="_blank" class="imCssLink">this page</a>. The code contains the implementation of CIRCLE, ready to use out-of-the-box, together with evaluation scripts and documentation.</div><div>If you find our work useful, please consider starring the repository!</div></div>]]></description>
			<pubDate>Wed, 25 Feb 2026 20:17:00 GMT</pubDate>
			<enclosure url="https://www.marcogarosi.it/blog/files/circle-code-available_thumb.webp" length="7977269" type="image/webp" />
			<link>https://www.marcogarosi.it/blog/?project-page-and-code-for-circle-are-available</link>
			<guid isPermaLink="false">https://www.marcogarosi.it/blog/rss/000000018</guid>
		</item>
		<item>
			<title><![CDATA[CIRCLE has been accepted to CVPR Findings]]></title>
			<author><![CDATA[Marco Garosi]]></author>
			<category domain="https://www.marcogarosi.it/blog/index.php?category=Research"><![CDATA[Research]]></category>
			<category>imblog</category>
			<description><![CDATA[<div id="imBlogPost_000000017"><div class="imHeading1">Exciting news: our latest research is headed to CVPR Findings 2026!</div><div><br></div><div>I am thrilled to announce that our latest paper, investigating the true classification potential of large multimodal models (LMMs), has been accepted to <b data-path-to-node="3" data-index-in-node="153">CVPR Findings 2026</b>!</div><div><br></div><div>For a long time, the computer vision community has relied on a generally accepted rule of thumb: use CLIP-like contrastive vision-language models (VLMs) for zero-shot classification, and save LMMs for more complex, generative tasks.</div><div>In our new paper, we challenge that status quo.</div><div><br></div><hr data-path-to-node="6"><div class="imHeading2">Rethinking multimodal classification</div><div>We wanted to know what happens when you leverage a frequently overlooked superpower of LMMs: <b data-path-to-node="8" data-index-in-node="93">in-context learning</b>. Here is a brief look at what we discovered:</div><ul data-path-to-node="9"><li><div><b data-path-to-node="9,0,0" data-index-in-node="0">Matching the specialists:</b> When we benchmarked state-of-the-art LMMs on diverse closed-world classification datasets, we found that providing just a few in-context examples allows LMMs to match—and sometimes surpass—contrastive VLMs using cache-based adapters.</div></li><li><div><b data-path-to-node="9,1,0" data-index-in-node="0">The open-world hurdle:</b> We then pushed LMMs into the much more challenging "open-world" setting. Because LMMs are generative, they naturally fit this task. However, they tend to stumble when provided with imperfect context information.</div></li><li><div><b data-path-to-node="9,2,0" data-index-in-node="0">Introducing CIRCLE:</b> To solve this, we developed <b data-path-to-node="9,2,0" data-index-in-node="48">CIRCLE</b>, a simple, completely <i data-path-to-node="9,2,0" data-index-in-node="77">training-free</i> method. CIRCLE assigns pseudo-labels to in-context examples and iteratively refines them using the available context itself.</div></li></ul><div class="imHeading2">A unified future</div><div>Through extensive experiments, we demonstrated that <b data-path-to-node="11" data-index-in-node="52">CIRCLE</b> establishes a highly robust baseline for open-world classification. It consistently outperforms VLM equivalents, proving that large multimodal models have the potential to act as powerful, unified classifiers and flexible alternatives to highly specialized models.</div><div>I can't wait to share more details with the community at CVPR 2026. Stay tuned for the full paper release and the accompanying code!</div></div>]]></description>
			<pubDate>Mon, 23 Feb 2026 20:06:00 GMT</pubDate>
			<enclosure url="https://www.marcogarosi.it/blog/files/circle-accepted_thumb.webp" length="9152014" type="image/webp" />
			<link>https://www.marcogarosi.it/blog/?circle-has-been-accepted-to-cvpr-findings</link>
			<guid isPermaLink="false">https://www.marcogarosi.it/blog/rss/000000017</guid>
		</item>
		<item>
			<title><![CDATA[New website for MHUG]]></title>
			<author><![CDATA[Marco Garosi]]></author>
			<category domain="https://www.marcogarosi.it/blog/index.php?category=Misc"><![CDATA[Misc]]></category>
			<category>imblog</category>
			<description><![CDATA[<div id="imBlogPost_000000016">I have worked on the new website of my research group MHUG (Multimedia and Human Understanding Group) at the University of Trento. <span class="fs14lh1-5">The new website is now live at </span><span class="fs14lh1-5"><a href="https://mhug.disi.unitn.it/" target="_blank" class="imCssLink">mhug.disi.unitn.it</a></span><span class="fs14lh1-5">.</span><div><span class="fs14lh1-5">Some cool new features I have introduced:</span></div><div><ul><li><span class="fs14lh1-5">Profile pages: now everyone has their own dynamic page, which lists contacts and published papers.</span></li><li><span class="fs14lh1-5">Paper linking: papers are now linked to their authors, so navigating through our research is easier than ever!</span></li><li><span class="fs14lh1-5">Open positions: want to join our lab? Take a look at the open positions or apply spontaneously.</span></li></ul><div><span class="fs14lh1-5"><br></span></div><div><span class="fs14lh1-5">The website is being continuously updated to accommodate all our needs and requirements.</span></div></div></div>]]></description>
			<pubDate>Fri, 28 Nov 2025 17:27:00 GMT</pubDate>
			<enclosure url="https://www.marcogarosi.it/blog/files/mhug_thumb.webp" length="1214215" type="image/webp" />
			<link>https://www.marcogarosi.it/blog/?new-website-for-mhug</link>
			<guid isPermaLink="false">https://www.marcogarosi.it/blog/rss/000000016</guid>
		</item>
		<item>
			<title><![CDATA[ComCa is available on arXiv]]></title>
			<author><![CDATA[Marco Garosi]]></author>
			<category domain="https://www.marcogarosi.it/blog/index.php?category=Research"><![CDATA[Research]]></category>
			<category>imblog</category>
			<description><![CDATA[<div id="imBlogPost_000000015"><div>Starting today, <a href="https://www.marcogarosi.it/comca.html" class="imCssLink" onclick="return x5engine.utils.location('https://www.marcogarosi.it/comca.html', null, false)">ComCa</a> is openly available on <a href="https://arxiv.org/abs/2503.19145" target="_blank" class="imCssLink">arXiv</a>, waiting to be read by all researchers in the community of open-vocabulary detection, working in training-free, retrieval-based settings, and specifically <span class="fs14lh1-5"><b>attribute detection</b></span>.</div></div>]]></description>
			<pubDate>Wed, 26 Mar 2025 06:59:00 GMT</pubDate>
			<enclosure url="https://www.marcogarosi.it/blog/files/large-936398_thumb.webp" length="194335" type="image/webp" />
			<link>https://www.marcogarosi.it/blog/?comca-is-available-on-arxiv</link>
			<guid isPermaLink="false">https://www.marcogarosi.it/blog/rss/000000015</guid>
		</item>
		<item>
			<title><![CDATA[Project page and code for ComCa are available]]></title>
			<author><![CDATA[Marco Garosi]]></author>
			<category domain="https://www.marcogarosi.it/blog/index.php?category=Research"><![CDATA[Research]]></category>
			<category>imblog</category>
			<description><![CDATA[<div id="imBlogPost_000000014">Code for my conference paper <a href="https://www.marcogarosi.it/comca.html" class="imCssLink" onclick="return x5engine.utils.location('https://www.marcogarosi.it/comca.html', null, false)">ComCa</a> is now openly available on <a href="https://github.com/marco-garosi/ComCa" target="_blank" class="imCssLink">GitHub</a>, along with the <a href="https://comca-attributes.github.io/" target="_blank" class="imCssLink">project page</a>.<div><br></div><div>The code has been developed with the latest libraries employed in deep learning, such as:</div><div><ul><li><a href="https://pytorch.org/" target="_blank" class="imCssLink">PyTorch</a>, as the main library to handle tensors, define and load models, and perform heavy computation;</li><li><a href="https://huggingface.co/docs/transformers/index" target="_blank" class="imCssLink">Transformers</a> by <a href="https://huggingface.co/" target="_blank" class="imCssLink">HuggingFace</a>, to load pre-trained 2D vision-language models such as CLIP and SigLIP;</li><li><a href="https://github.com/facebookresearch/faiss" target="_blank" class="imCssLink">Faiss</a> by Meta's <a href="https://ai.facebook.com/" target="_blank" class="imCssLink">Fundamental AI Research</a> group to efficiently search in large collections of dense vectors. Specifically, the library is used to retrieve data from web-scale databases with millions of elements in an efficient way.</li></ul><div><br></div></div><div>Code is available <a href="https://github.com/marco-garosi/ComCa" target="_blank" class="imCssLink">here</a>.</div></div>]]></description>
			<pubDate>Mon, 24 Mar 2025 07:00:00 GMT</pubDate>
			<enclosure url="https://www.marcogarosi.it/blog/files/large-1283624_thumb.webp" length="224321" type="image/webp" />
			<link>https://www.marcogarosi.it/blog/?project-page-and-code-for-comca-are-available</link>
			<guid isPermaLink="false">https://www.marcogarosi.it/blog/rss/000000014</guid>
		</item>
		<item>
			<title><![CDATA[ComCa has been accepted to CVPR]]></title>
			<author><![CDATA[Marco Garosi]]></author>
			<category domain="https://www.marcogarosi.it/blog/index.php?category=Research"><![CDATA[Research]]></category>
			<category>imblog</category>
			<description><![CDATA[<div id="imBlogPost_000000013">My paper ComCa, short for "Compositional Caching for Training-free Open-vocabulary Attribute Detection", has been accepted to the <a href="https://cvpr.thecvf.com/Conferences/2025" target="_blank" class="imCssLink">CVPR 2025</a> <a href="https://www.marcogarosi.it/conferences.html" class="imCssLink" onclick="return x5engine.utils.location('https://www.marcogarosi.it/conferences.html', null, false)">conference</a>! 🎊<div>This research leverages pre-trained vision-language models (VLMs) for the task of <b>attribute detection</b>, which involves identifying all attributes of an object, such as color, material and texture.</div><div>ComCa, our training-free method, surpasses all other training-free models and is even competitive with training-based approaches.</div><div>By introducing a <b>caching strategy</b> that stores examples of objects with specific attributes, ComCa can refine the VLM's predictions at test time requiring no specific training nor manually annotated data. Indeed, the cache is constructed automatically by leveraging web-scale databases, such as CC12M, and large language models (LLMs) such as GPT by OpenAI.</div><div><br></div><div>Code will be released shortly, and the camera-ready version of the paper will be uploaded on arXiv as soon as it is ready.</div></div>]]></description>
			<pubDate>Thu, 27 Feb 2025 14:51:00 GMT</pubDate>
			<enclosure url="https://www.marcogarosi.it/blog/files/cvpr2025_nashville_tn_thumb.webp" length="637612" type="image/webp" />
			<link>https://www.marcogarosi.it/blog/?comca-has-been-accepted-to-cvpr</link>
			<guid isPermaLink="false">https://www.marcogarosi.it/blog/rss/000000013</guid>
		</item>
		<item>
			<title><![CDATA[COPS is available on arXiv]]></title>
			<author><![CDATA[Marco Garosi]]></author>
			<category domain="https://www.marcogarosi.it/blog/index.php?category=Research"><![CDATA[Research]]></category>
			<category>imblog</category>
			<description><![CDATA[<div id="imBlogPost_00000000F">Starting today, <span class="fs14lh1-5"><a href="https://www.marcogarosi.it/cops.html" class="imCssLink" onclick="return x5engine.utils.location('https://www.marcogarosi.it/cops.html', null, false)">COPS</a> 👮‍♂️is openly available on <a href="https://arxiv.org/abs/2412.04247" target="_blank" class="imCssLink">arXiv</a>, waiting to be read by all researchers in the community of 3D data processing, 3D point cloud understanding, and specifically <b>3D point cloud part segmentation</b>.</span></div>]]></description>
			<pubDate>Fri, 06 Dec 2024 07:32:00 GMT</pubDate>
			<enclosure url="https://www.marcogarosi.it/blog/files/large-5499156_thumb.webp" length="291598" type="image/webp" />
			<link>https://www.marcogarosi.it/blog/?cops-is-available-on-arxiv</link>
			<guid isPermaLink="false">https://www.marcogarosi.it/blog/rss/00000000F</guid>
		</item>
		<item>
			<title><![CDATA[Code for COPS is available]]></title>
			<author><![CDATA[Marco Garosi]]></author>
			<category domain="https://www.marcogarosi.it/blog/index.php?category=Research"><![CDATA[Research]]></category>
			<category>imblog</category>
			<description><![CDATA[<div id="imBlogPost_000000010">Code for my conference paper <span class="fs14lh1-5"><a href="https://www.marcogarosi.it/cops.html" class="imCssLink" onclick="return x5engine.utils.location('https://www.marcogarosi.it/cops.html', null, false)">COPS</a> 👮‍♂️is openly available on <a href="https://github.com/marco-garosi/COPS" target="_blank" class="imCssLink">GitHub</a>, along with the <a href="https://3d-cops.github.io/" target="_blank" class="imCssLink">project page</a>.</span><div><span class="fs14lh1-5"><br></span></div><div><span class="fs14lh1-5">The code has been developed with the latest technologies used for deep learning. Notably:</span></div><div><ul><li><span class="fs14lh1-5"><a href="https://pytorch.org/" target="_blank" class="imCssLink">PyTorch</a>, as the main library to handle vectors, tensors, and their mathematical manipulation;</span></li><li><span class="fs14lh1-5"><a href="https://pytorch3d.org/" target="_blank" class="imCssLink">PyTorch3D</a>, to handle 3D data efficiently;</span></li><li><span class="fs14lh1-5"><a href="https://pytorch-geometric.readthedocs.io/en/latest/" target="_blank" class="imCssLink">PyTorch Geometric</a>, to access some 3D point cloud part segmentation datasets and to easily manipulate large point clouds;</span></li><li><span class="fs14lh1-5"><a href="https://huggingface.co/docs/transformers/index" target="_blank" class="imCssLink">Transformers library </a>by <a href="https://huggingface.co/" target="_blank" class="imCssLink">HuggingFace</a>, to load pre-trained 2D vision foundation models.</span></li></ul><div><span class="fs14lh1-5"><br></span></div><div><span class="fs14lh1-5">Code is available <a href="https://github.com/marco-garosi/COPS" target="_blank" class="imCssLink">here</a>.</span></div></div></div>]]></description>
			<pubDate>Fri, 29 Nov 2024 07:38:00 GMT</pubDate>
			<enclosure url="https://www.marcogarosi.it/blog/files/large-1841550_thumb.webp" length="222519" type="image/webp" />
			<link>https://www.marcogarosi.it/blog/?code-for-cops-is-available</link>
			<guid isPermaLink="false">https://www.marcogarosi.it/blog/rss/000000010</guid>
		</item>
		<item>
			<title><![CDATA[COPS has been accepted to WACV]]></title>
			<author><![CDATA[Marco Garosi]]></author>
			<category domain="https://www.marcogarosi.it/blog/index.php?category=Research"><![CDATA[Research]]></category>
			<category>imblog</category>
			<description><![CDATA[<div id="imBlogPost_00000000E">My paper <span class="fs14lh1-5"><a href="https://www.marcogarosi.it/cops.html" class="imCssLink" onclick="return x5engine.utils.location('https://www.marcogarosi.it/cops.html', null, false)">COPS </a>👮‍♂️ </span><span class="fs14lh1-5">has been accepted to the prestigious <a href="https://wacv2025.thecvf.com/" target="_blank" class="imCssLink">WACV 2025</a> <a href="https://www.marcogarosi.it/conferences.html" class="imCssLink" onclick="return x5engine.utils.location('https://www.marcogarosi.it/conferences.html', null, false)">conference</a>! 🎉</span><div><span class="fs14lh1-5">This research dives into the innovative realm of <b>3D point cloud part segmentation</b>, utilizing plain 2D vision foundation models to achieve remarkable results in a zero-shot, training-free setting.</span></div><div><span class="fs14lh1-5">By bridging the gap between 2D and 3D data processing, and introducing a novel <b>geometric feature aggregation</b> (GFA) module to incorporate geometrical knowledge into the vision model's features, COPS not only enhances the accuracy of part segmentation but also paves the way for more efficient and scalable applications in various fields, including robotics and autonomous vehicles.</span></div><div><br></div><div>If you are interested in COPS, <a href="https://www.marcogarosi.it/cops.html" class="imCssLink" onclick="return x5engine.utils.location('https://www.marcogarosi.it/cops.html', null, false)">take a look at it</a>!</div></div>]]></description>
			<pubDate>Tue, 29 Oct 2024 07:24:00 GMT</pubDate>
			<enclosure url="https://www.marcogarosi.it/blog/files/large-1732395_thumb.webp" length="226019" type="image/webp" />
			<link>https://www.marcogarosi.it/blog/?cops-has-been-accepted-to-wacv</link>
			<guid isPermaLink="false">https://www.marcogarosi.it/blog/rss/00000000E</guid>
		</item>
	</channel>
</rss>