Senior Data Scientist (OCR/CV)
Summary
Design, build, deploy, and improve ML systems for intelligent document understanding (OCR/CV) using Python, computer vision models, and LLMs in production at Workato's iPaaS platform.
About Workato
Workato is the leading Control and Execution Platform for Enterprise AI — the neutral platform enterprises trust to put AI to work across their business. Workato unifies data, applications, and processes into a single platform so AI can reliably orchestrate business processes in production at enterprise scale. Built on more than a decade of running mission-critical processes for over half the Fortune 500 — including Nasdaq, Amazon, Cisco, Vodafone, Atlassian, and Lucid Motors — Workato turns over 14,000 enterprise systems AI needs to act on into one governed execution layer. For more information, visit .
Why join us?
Ultimately, Workato believes in fostering a flexible, trust-oriented culture that empowers everyone to take full ownership of their roles. We are driven by innovation and looking for team players who want to actively build our company.
But, we also believe in balancing productivity with self-care. That’s why we offer all of our employees a vibrant and dynamic work environment along with a multitude of benefits they can enjoy inside and outside of their work lives.
If this sounds right up your alley, please submit an application. We look forward to getting to know you!
Also, feel free to check out why:
-
Business Insider named us an “enterprise startup to bet your career on”
-
Forbes’ Cloud 100 recognized us as one of the top 100 private cloud companies in the world
-
Deloitte Tech Fast 500 ranked us as the 17th fastest growing tech company in the Bay Area, and 96th in North America
-
Quartz ranked us the #1 best company for remote workers
Responsibilities
We are looking for an exceptional Senior Data Scientist (OCR/CV) to join our growing AI team and work on OCR-focused product capabilities. You will design, build, deploy, and improve machine learning systems for intelligent document understanding. The project involves extracting information from complex document layouts, including tables, embedded images, multi-column text, forms, and other non-trivial visual elements, while maintaining high accuracy and robustness in production environments. This role is ideal for someone with strong analytical thinking, solid computer vision expertise, and hands-on experience building ML systems for real-world document understanding tasks. Experience with OCR pipelines and LLM-based post-processing or document understanding systems is highly desirable.
In this role, you will also be responsible to:
-
Build and improve AI services using LLMs and custom machine learning models for production use cases.
-
Design, develop, and operate ML/LLM systems end-to-end, from prototyping to deployment and monitoring.
-
Write high-quality Python code that is testable, maintainable, and efficient.
-
Improve validation, observability, and performance monitoring for ML services (quality, latency, reliability, cost).
-
Partner cross-functionally with product managers, platform engineers, and other stakeholders to ship AI-powered product capabilities.
-
Evaluate and improve existing implementations by identifying bottlenecks, bugs, and opportunities for optimization.
-
Design controlled experiments to test the features for our AI-based products and perform deep analysis from the results to find actionable insights
-
Contribute to technical design and code reviews, helping raise engineering quality across the team.
-
Experiment and iterate on model behavior, prompting, retrieval, tool use, or orchestration strategies to improve user outcomes.
Required Qualifications
-
Bachelor’s or Master’s degree in Computer Science, Engineering, Mathematics, Statistics, or equivalent practical experience
-
3+ years of experience in Machine Learning Engineering, Data Science, or a similar role.
-
Strong Python programming skills.
-
Hands-on experience with CV and/or LLM-based systems.
-
Experience deploying and operating ML services in production.
-
Strong understanding of software engineering fundamentals (testing, code quality, debugging, version control).
-
Ability to work collaboratively in a fast-moving environment and drive projects with ownership.
Preferred Qualifications
-
Experience with tool-use agents or workflow-aware AI systems.
-
Experience building AI products in enterprise SaaS environments.
-
Experience with A/B testing and statistical significance techniques.
-
Experience with LLMOps/MLOps tooling and practices (monitoring, evaluation pipelines, model rollout, CI/CD).
-
Experience working with modern data warehouses such as Amazon Redshift or Snowflake.
Job Req ID: 2455
Skills
As published by greenhouse · 13 questions · 3 written answers
Basics
First Name, Last Name, Email, Phone, Resume/CV, Cover Letter
Short answers (6)
- LinkedIn Profile
- Years of experience building AI products in enterprise SaaS environments?
- Years of experience with tool-use agents or workflow-aware AI systems?
- Years of experience with LLMOps/MLOps tooling and practices (monitoring, evaluation pipelines, model rollout, CI/CD)?
- What is your experience with LLMOps/MLOps, including tools used for monitoring, evaluation, and model rollout?
- What is your expected compensation?
Pick from a list (4)
- Residency Status in Singapore
- Which modern data warehouses have you worked with hands-on?
- Which OCR, Computer Vision, or document-understanding tools/frameworks have you used?
- Have you used LLMs to improve OCR/document extraction outputs—for example, validation, structured data extraction, post-processing, or error correction?
Written answers (3)
- Summarize your experience with tool-use agents or workflow-aware AI systems, including specific frameworks or platforms used.
- Have you worked on extracting information from complex documents such as tables, forms, multi-column text, scanned PDFs, or embedded images? Please share a brief example.
- What is your level of proficiency in Python, and which ML/CV libraries have you used most frequently?