ABBYY

Senior Machine Learning Engineer, Synthetic Data & Document Understanding

Senior · Гибрид · Бангалор, Индия · Английский B2

Навыки

  • Computer Vision
  • Решения на данных
  • Data Quality
  • Лидерство
  • Machine Learning
  • NLP
  • Zapier / Make / n8n
Ещё 4
  • Ответственность за результат
  • Python
  • PyTorch
  • Роадмап

О компании и продукте

  • We are seeking a Senior Machine Learning Engineer – Synthetic Data & Document Understanding to own the synthetic data generation track within ABBYY’s Document AI Data team .
  • This role focuses on building generative pipelines that produce high-quality, diverse, and realistic synthetic training data at scale. You will ensure synthetic data meaningfully improves downstream model performance by maintaining strong alignment with real-world document structures, formats, and statistical properties.
  • This is an ideal role for engineers who combine deep generative modeling expertise with rigorous data quality evaluation and production engineering skills .
  • Join ABBYY and be part of a team that celebrates your unique work style. With flexible work options, a supportive team, and rewards that reflect your value, you can focus on what matters most – driving your growth, while fueling ours.

Задачи

  • Technical Development & Innovation
  • Design and implement pipelines that analyze real documents to inform high-fidelity synthetic data generation
  • Build generative systems capable of producing documents across diverse formats, layouts, and domains
  • Develop evaluation frameworks to ensure synthetic data maintains distributional fidelity and diversity
  • Research and apply generative modeling techniques suited for document AI training
  • Identify and mitigate quality issues to ensure synthetic data is effective for downstream model training
  • Partner with Modeling teams to measure the impact of synthetic data on model performance
  • Project Ownership & Leadership
  • Own the synthetic data generation track end-to-end , from architecture to quality validation
  • Drive architectural decisions balancing quality, diversity, scale, and cost efficiency
  • Define and maintain data quality metrics and generation dashboards
  • Collaborate closely with annotation teams to ensure compatibility with downstream pipelines
  • Contribute to roadmap planning alongside Principal-level leadership
  • Infrastructure & Scale
  • Build scalable pipelines capable of generating millions of synthetic training examples
  • Implement post-processing, filtering, and validation mechanisms to remove low-quality outputs
  • Design cost-efficient workflows balancing compute, quality, and throughput
  • Develop monitoring systems to detect distribution shifts or quality degradation over time
  • Collaborate with Platform teams on compute orchestration, storage, and scheduling

Требования

  • Education & Experience
  • MS or PhD in Computer Science, Engineering, Mathematics, or related field
  • 5+ years of experience in Machine Learning / AI , with focus on
  • Generative models
  • Vision-Language Models (VLMs)
  • Synthetic data systems
  • Proven experience building and evaluating synthetic data pipelines for ML training
  • Strong background in data quality evaluation and statistical analysis
  • Technical Expertise
  • Deep expertise in Vision-Language Models and document understanding (layout, structure, semantics)
  • Strong knowledge of generative modeling for structured and semi-structured data
  • Understanding of what makes synthetic data valuable
  • Distributional fidelity
  • Diversity
  • Realistic noise patterns
  • Domain coverage
  • Strong programming skills in Python with experience in PyTorch or similar frameworks
  • Experience evaluating data quality via automated metrics and downstream model impact
  • Familiarity with large-scale data pipelines, cloud environments, and experiment tracking
  • Leadership & Communication
  • Proven ability to independently own complex technical workstreams
  • Strong collaboration across data, modeling, and platform teams
  • Ability to clearly communicate data quality and generation trade-offs
  • Data-driven mindset with strong attention to coverage gaps and quality signals
  • Here are some of our local benefits

Паспорт вакансии

История публикации

Появилась в Вакандии26 дней
Перепубликациинетпубликовалась один раз
Проверяли на источникеВидели 26 дней назад
Среди похожихНет данных166 из 30 · у похожих вакансий почти одинаковый возраст — сравнивать нечего

Откуда что взялось

Отмечено то, что вывели мы. Без пометки — значение назвал работодатель.

ГрейдSeniorвычитано из текста вакансии
Формат работыГибрид
ГеографияБангалор, Индиявычитано из текста вакансии
Зарплата≈ 19 058 USD в месяцнаша оценка, в вакансии не названа

Почему на этом месте в выдаче

Порядок выдачи объявлен контрактом: свежесть решает между днями, полнота и зарплата — внутри дня.

Полнота карточки1004 из 4 полей: грейд, формат, география, зарплата
Зарплата названа0вилки работодателя нет, показана наша оценка

Проверка Вакандии

Источники и свежесть

Тип источника
Карьерный сайт работодателя
Найдено публикаций
1
Посмотреть публикации и даты
  • greenhouseОсновная публикация · 2026-05-21

Работодатель

ABBYY

12 активных вакансий · вилка работодателя указана в 0%

Открыть профиль компании

Безопасность

Отклик уходит на сайт источника

Вакандия показывает вакансию, но не отправляет отклик и не проверяет работодателя. Сам отклик вы оставляете на внешнем сайтеjob-boards.eu.greenhouse.io.

Признаки мошенничества
  • Просят предоплату, «залог» или деньги за обучение и оборудование.
  • Требуют код из SMS, данные банковской карты или доступ к «Госуслугам».
  • Быстро уводят в мессенджер и торопят с решением.
  • Обещают большой доход без опыта и без деталей задач.

Настоящий работодатель не просит денег и платёжных данных до трудоустройства.

Продолжить поиск

Похожие вакансии

Причина сходства указана на каждой карточке