We collect paired sight and olfaction by walking the city and recording whatever is there to be smelled — indoors and out, over sixty sessions and thirty-six locations!
A Large MultimodalDataset for Olfaction
While olfaction is central to how animals perceive the world, this rich chemical sensory modality remains largely inaccessible to machines. One key bottleneck is the lack of diverse, multimodal olfactory data collected in natural settings. We present New York Smells, a large-scale dataset of paired image and olfactory signals captured in-the-wild. Our dataset contains 7,000 smell-image pairs from 3,500 distinct objects across diverse indoor and outdoor environments, and it is 70× larger than prior olfactory datasets. We create benchmarks for three tasks: cross-modal smell-to-image retrieval, recognizing scenes, objects, and materials from smell alone, and fine-grained discrimination between grass species. Models trained on raw olfactory signals outperform widely-used hand-crafted features, and visual data enables learning of olfactory representations.
We collect paired sight and olfaction by walking the city and recording whatever is there to be smelled — indoors and out, over sixty sessions and thirty-six locations!
54 materials · 86 objects · 36 places · 6,934 samples

We walk through New York City and capture paired olfaction and visual signals using a camera mounted to a Cyranose 320 electronic nose on a custom 3D-printed sensor rig. We also capture depth, temperature, humidity, and ambient VOC concentrations.
We train a linear probe on our embedding space to classify objects, materials, and scenes from smell alone. Images are shown for visualization purposes only — hover a row to uncover its photograph.
We show cross-modal retrieval given a smell signal as input and use our joint embedding space to retrieve images.
We train general-purpose olfactory representations by contrasting smell against vision. A vision encoder and an olfaction encoder embed a batch into one space; the objective pulls each co-occurring pair together and pushes every mismatched pair apart.

View of the Upper West Side skyline from Central Park, New York City.
@article{ozguroglu2025smell,
title={New York Smells: A Large Multimodal Dataset for Olfaction},
author={Ozguroglu, Ege and Liang, Junbang and Liu, Ruoshi and Chiquier, Mia and DeTienne, Michael and Qian, Wesley Wei and Horowitz, Alexandra and Owens, Andrew and Vondrick, Carl},
journal={arXiv preprint arXiv:2511.20544},
year={2025}
}