Direct Comparison with Flamingo and OpenFlamingo
Hi there,
Congratulations on this great success.
I noticed that in the Model Card it says, "We note that since IDEFICS was trained on PMD (which contains COCO), the evaluation numbers on COCO are not directly comparable with Flamingo and OpenFlamingo since they did not explicitly have this dataset in the training mixture."
However, as far as I know, datasets like VQAv2 and OKVQA also build on images from COCO. Are IDEFICS's results directly comparable with Flamingo and OpenFlamingo on these benchmarks as well?
Thanks.
Hi, thanks for your question!
It's not a straightforward question.
We argue that it is comparable in the "this is not the same task" sense. COCO (as an image captioning task) was part of the training and evaluation suite. however, VQA was not part of the evaluation suite.
Although some of the images might be in the training, it is also very unlikely that any of the qa samples would be in the training text verbatim, which makes it questionable whether there is leakage.