[VL-JEPA] Joint Embedding Predictive Architecture for Vision-Language. V-JEPA Vision Language Models