Arrastra un PDF aquí, pega una URL o toca para seleccionar un archivo
Reka
Singapore, UK, USA
In this role, you’ll work closely with model researchers, data infrastructure engineers, and cross-functional partners to make sure our data is high quality and can be produced at petabyte scale in a reliable, efficient way. From understanding how data choices show up in model behavior, to building processing pipelines and running the compute behind them, you’ll help ensure our models are trained on the best data we can get. What you’ll do Work with model researchers to define what “good data” means for our models, including quality metrics, validation checks, and acceptance thresholds Explore open source datasets and create internal ones most suitable to build fundamental World Models Build algorithms for automated data quality assessment, data domain mixtures, and domain adaptation from synthetic to real data. Track datasets, metadata, provenance, and versions so experiments are reproducible and it’s clear what data went into which training and evaluation runs Own CI/CD and development tooling for the data stack (GitHub, Python, PyTorch), and automate repetitive workflows to reduce friction Track and optimize throughput, storage, and compute utilization across pipelines and relat
Strong ML and deep learning fundamentals with experience building and operating large-scale data and/or compute systems Comfortable moving between research questions and production engineering: you can dig into data, run analyses, and
Ref. 21YF5
Arrastra un PDF aquí, pega una URL o toca para seleccionar un archivo
Tu Anuncio para Redes Sociales
×Esta imagen está optimizada para Instagram y LinkedIn. ¡Compártela para atraer talento!
Se envía un enlace que abre esta oferta directamente.
Tu Anuncio para TikTok / Reels
Esta imagen vertical está optimizada para TikTok, Instagram Reels y YouTube Shorts.