Remote Native-Language AI Benchmark Engineer
Guarda esta oferta y sigue tu búsqueda
Crea una cuenta gratis para guardar empleos, crear alertas y volver a esta oferta desde tu panel.
Al continuar, aceptas nuestros Términos & Política de Privacidad.
LILT is seeking experienced software engineers to design, build, and validate benchmarks for multilingual models. This remote, freelance role focuses on creating high-quality tasks in the candidate’s native language, evaluating agents, and writing deterministic verifier scripts.
You’ll work across data, prompts, and evaluation rubrics, collaborating with a global linguistics and ML team. Ideal candidates have 5+ years of software engineering experience, strong Python and shell skills, and a deep