Senior Machine Learning Engineer, Synthetic Data & Document Understanding
Senior · Гибрид · Бангалор, Индия · Английский B2
Навыки
Computer Vision
Решения на данных
Data Quality
Лидерство
Machine Learning
NLP
Zapier / Make / n8n
Ещё 4
Ответственность за результат
Python
PyTorch
Роадмап
О компании и продукте
We are seeking a Senior Machine Learning Engineer – Synthetic Data & Document Understanding to own the synthetic data generation track within ABBYY’s Document AI Data team .
This role focuses on building generative pipelines that produce high-quality, diverse, and realistic synthetic training data at scale. You will ensure synthetic data meaningfully improves downstream model performance by maintaining strong alignment with real-world document structures, formats, and statistical properties.
This is an ideal role for engineers who combine deep generative modeling expertise with rigorous data quality evaluation and production engineering skills .
Join ABBYY and be part of a team that celebrates your unique work style. With flexible work options, a supportive team, and rewards that reflect your value, you can focus on what matters most – driving your growth, while fueling ours.
Задачи
Technical Development & Innovation
Design and implement pipelines that analyze real documents to inform high-fidelity synthetic data generation
Build generative systems capable of producing documents across diverse formats, layouts, and domains
Develop evaluation frameworks to ensure synthetic data maintains distributional fidelity and diversity
Research and apply generative modeling techniques suited for document AI training
Identify and mitigate quality issues to ensure synthetic data is effective for downstream model training
Partner with Modeling teams to measure the impact of synthetic data on model performance
Project Ownership & Leadership
Own the synthetic data generation track end-to-end , from architecture to quality validation
Drive architectural decisions balancing quality, diversity, scale, and cost efficiency
Define and maintain data quality metrics and generation dashboards
Collaborate closely with annotation teams to ensure compatibility with downstream pipelines
Contribute to roadmap planning alongside Principal-level leadership
Infrastructure & Scale
Build scalable pipelines capable of generating millions of synthetic training examples
Implement post-processing, filtering, and validation mechanisms to remove low-quality outputs
Design cost-efficient workflows balancing compute, quality, and throughput
Develop monitoring systems to detect distribution shifts or quality degradation over time
Collaborate with Platform teams on compute orchestration, storage, and scheduling
Требования
Education & Experience
MS or PhD in Computer Science, Engineering, Mathematics, or related field
5+ years of experience in Machine Learning / AI , with focus on
Generative models
Vision-Language Models (VLMs)
Synthetic data systems
Proven experience building and evaluating synthetic data pipelines for ML training
Strong background in data quality evaluation and statistical analysis
Technical Expertise
Deep expertise in Vision-Language Models and document understanding (layout, structure, semantics)
Strong knowledge of generative modeling for structured and semi-structured data
Understanding of what makes synthetic data valuable
Distributional fidelity
Diversity
Realistic noise patterns
Domain coverage
Strong programming skills in Python with experience in PyTorch or similar frameworks
Experience evaluating data quality via automated metrics and downstream model impact
Familiarity with large-scale data pipelines, cloud environments, and experiment tracking
Leadership & Communication
Proven ability to independently own complex technical workstreams
Strong collaboration across data, modeling, and platform teams
Ability to clearly communicate data quality and generation trade-offs
Data-driven mindset with strong attention to coverage gaps and quality signals
Here are some of our local benefits
Паспорт вакансии
История публикации
Появилась в Вакандии26 дней
Перепубликациинетпубликовалась один раз
Проверяли на источникеВидели 26 дней назад
Среди похожихНет данных166 из 30 · у похожих вакансий почти одинаковый возраст — сравнивать нечего
Откуда что взялось
Отмечено то, что вывели мы. Без пометки — значение назвал работодатель.
ГрейдSeniorвычитано из текста вакансии
Формат работыГибрид
ГеографияБангалор, Индиявычитано из текста вакансии
Зарплата≈ 19 058 USD в месяцнаша оценка, в вакансии не названа
Почему на этом месте в выдаче
Порядок выдачи объявлен контрактом: свежесть решает между днями, полнота и зарплата — внутри дня.
Полнота карточки1004 из 4 полей: грейд, формат, география, зарплата
Зарплата названа0вилки работодателя нет, показана наша оценка
Проверка Вакандии
Источники и свежесть
Тип источника
Карьерный сайт работодателя
Найдено публикаций
1
Посмотреть публикации и даты
greenhouseОсновная публикация · 2026-05-21
A
Работодатель
ABBYY
12 активных вакансий · вилка работодателя указана в 0%
Вакандия показывает вакансию, но не отправляет отклик и не проверяет работодателя. Сам отклик вы оставляете на внешнем сайте — job-boards.eu.greenhouse.io.
Признаки мошенничества
Просят предоплату, «залог» или деньги за обучение и оборудование.
Требуют код из SMS, данные банковской карты или доступ к «Госуслугам».
Быстро уводят в мессенджер и торопят с решением.
Обещают большой доход без опыта и без деталей задач.
Настоящий работодатель не просит денег и платёжных данных до трудоустройства.
Продолжить поиск
Похожие вакансии
Причина сходства указана на каждой карточке
SL
Почему похожа: похожая специализация · тот же грейд