how to generate most matching text set

thanks for sharing your great work.

I am confused about this step
"we select the most matching caption pairs from the dataset of each image v to form an augmented caption set t ={t1, t2, ..., tM }"

because the dataset only gives single image-text pair, how could you find multiple matched texts for an single image