Job description
As a key contributor, you will design and train innovative Vision-Language Models aimed at analyzing intricate construction site data while steering the research agenda for spatial intelligence. Your responsibilities will also involve developing scalable training pipelines and enhancing inference efficiency to manage extensive video data. Applicants should possess over 6 years of practical experience in deep learning and transformer architectures, demonstrating advanced skills in Python. A degree in Computer Science, Machine Learning, or a relevant discipline is essential, paired with familiarity with prominent deep learning frameworks.


