Kastan Day
Results
2
repositories owned by
Kastan Day
video-pretrained-transformer
50
Stars
8
Forks
Watchers
Multi-model video-to-text by combining embeddings from Flan-T5 + CLIP + Whisper + SceneGraph. The 'backbone LLM' is pre-trained from scratch on YouTube (YT-1B dataset).