In the first year of her doctoral research, Abigail Grassick collected hours of underwater video as she studied fish communities living on bommies – individual clumps of coral – on the seafloor off the coast of Curaçao.
A doctoral candidate in the field of computational biology, she tried using computer vision models to automatically identify behaviors of certain fish species. But the existing models totally failed – they couldn’t even keep tabs on an individual fish, let alone identify when they were feeding or trying to evade a predator.
To help pave the way for better models, Grassick and a team of Cornell researchers and colleagues created WildFin, a data set composed of nine hours of video that they labeled frame-by-frame with fish behaviors. They hope computer scientists will use the video to develop and test models that better interpret fish behavior, and they invite fellow ecologists to release their own field videos to assist in the effort.
Grassick presented their work, “WildFin: An In-the-Wild Dataset for Fish Behavioral Recognition,” at the European Conference on Computer Vision on Sept. 8 in Malmö, Sweden.
Existing models perform poorly because they are trained on very little wildlife footage, said Andrew Hein, associate professor of computational biology in the College of Agriculture and Life Sciences and co-author of the study. “People are training these computer vision models on internet-scale data. Well, what fraction of internet-scale data are scientific imagery? The answer is, it’s a tiny fraction.”
Better computer vision models would not only allow ecologists to answer their questions more efficiently, but would also enable researchers to use citizen science videos and crowdsourced data from naturalist platforms to monitor the health of ecosystems and log occurrences of rare species.
“My hope is that we really can unlock the value in citizen science footage,” Hein said, “and allow biologists who are going out into all these amazing places around the world to vastly increase the amount of data that we can acquire through these really expensive and time-consuming expeditions.”
Hein and Grassick collaborated with Jennifer Sun, assistant professor of computer science in the Cornell Ann S. Bowers College of Computing and Information Science. Sun had successfully used computer vision models to interpret behaviors of animals in the lab. But the wild-caught footage, with its roving bands of fish, variable lighting and shifting camera angles, proved to be a bigger challenge.
“When I first talked to Andrew, I never anticipated how hard his videos were to analyze,” Sun said. “He has hundreds of fish, and it’s hard to know if it’s even the same fish in subsequent frames.”
Grassick and a team of undergrads annotated the bommie videos by choosing from a list of behaviors to catalogue what the fish were doing in each frame – such as foraging, feeding or fighting. They also included footage where a diver tracked a single fish across the seafloor – similar to videos taken by divers on vacation – to see how models would handle the different types of footage. These videos were collected and annotated by colleagues at the University of Colorado, Boulder. All videos came from different sites off the coast of Curaçao.
Creating WildFin was incredibly time-consuming, Grassick said. Researchers spent about 1,400 hours in the field and 600 hours annotating, which ultimately yielded just nine hours of footage.
The researchers attempted to use this footage to fine-tune several standard computer vision models to classify fish behavior, but their performance was still poor.
Now that researchers have identified this gap in computer vision technology, they hope experts will use this dataset and analysis as a starting point to create improved models for studying fish behavior.
“It’s a multidisciplinary problem that needs to be addressed,” Grassick said. “You need statistics people, machine-learning people and ecologists all working together to figure out how to make this process less expensive, and how to still answer our questions.”
They also invite ecologists everywhere to unearth videos from their field work – much of it sitting on hard drives on lab shelves – to provide training data for new models. Just cleaning up that data is a challenge in itself, so Sun envisions that artificial intelligence agents could convert them into a useful format.
“If those datasets could be automatically ingested once humans are done analyzing them, in some sense, they can become alive again,” Sun said.
Co-authors on the paper include undergraduates Jerome Tze-Hou Hsu ’27, Ethan Lin ’26 and Max Whitton ’25; Ziang Liu and Haozheng Yu, both doctoral students in the field of computer science; Madelyn Hair, Liam Gutierrez and Michael Gil of the University of Colorado, Boulder; and Kristin Branson and Vivek Jayaraman of the Howard Hughes Medical Institute.
The research received partial support from the National Science Foundation and support for hardware through NVIDIA’s Academic Grant Program.
Patricia Waldron is a writer for the Cornell Ann S. Bowers College of Computing and Information Science.
