ML Operations Engineer (AI/LLM) - Mercari
Summary
ML Operations Engineer owning production ML/LLM model serving, deployment, and operations in Mercari's cloud-native environment using Kubernetes, Terraform, Python, and NVIDIA/TPU stacks (Triton, TensorRT-LLM, JAX).
本ポジションは日本語JDの用意がありません。
ML Operations Engineer (AI/LLM) - Mercari
- Employment Status: Full-time
- Work Hours: Full Flextime (no core time)
- Office: Roppongi
For more details, see the Overview of Our Positions section on our Careers site.
About Mercari
Circulate all forms of value to unleash the potential in all people
"What can I do to help society thrive with the finite resources we have?" The Mercari marketplace app was born in 2013 out of this thought by our founder Shintaro Yamada as he traveled the world. We believe that by circulating all forms of value, not just physical things and money, we can create opportunities for anyone to realize their dreams and contribute to society and the people around them. Mercari aims to use technology to connect people all over the world and create a world where anyone can unleash their potential. For more information about Mercari Group’s mission, see Mercari’s Culture Doc
Organization/Team Mission
Mercari Engineering Principles
Mercari Engineering Principles are a shared understanding that serves as the foundation of engineering beliefs and behavior at Mercari. The Engineering Principles are designed to complement the organizational identity (Mercari’s mission, values, and culture) from an engineering viewpoint.
These principles ultimately help us achieve Mercari’s mission by defining the ideal state we seek to realize in the long term.
- Passion For The Product
- Grow Together
- Solve Through Mechanisms
- Collaborate Openly
For more details, please see the following link:
- Engineering Culture
The AI / LLM Team’s mission is focused on three core pillars, “product”, "enablement" and "research", delivering new AI-driven features and user experiences to maximize product-facing impact for Mercari's business. We do this both through independent initiatives owned by our team, as well as by horizontally collaborating with product, engineering, and research teams across the entire organization.
As an MLOps engineer on the AI/LLM team, you will own how our machine learning and LLM models reach production and stay healthy there in our cloud-native environment. Your focus is the production serving, deployment, and operations that turn models into reliable, cost-efficient services, seamlessly integrating with our machine learning operations to serve tens of millions of users.
Work Responsibilities
- Data and Model Orchestration: Own the end-to-end orchestration of model inference, including integrating with DataServ for retrieval (BigQuery, BigTable, Valkey) and managing the Model Inference Gateway and Console.
- Model Serving and Deployment: Own production model serving on the cloud-native NVIDIA and TPU stacks (Triton Inference Server, TensorRT-LLM, JAX/TPU Gateways). Manage model repositories, dynamic batching, and concurrent model execution. Build CI/CD, rollout, and rollback paths for safe, scalable model deployment, and automate provisioning and lifecycle management (Terraform, Kubernetes) so the platform scales seamlessly across teams.
- Inference Performance Optimization: Profile and optimize deployments across LLM and non-LLM workloads, including model compilation, quantization, and batching strategies to hit latency and throughput targets while managing cost. Maintain performance baselines and regression detection.
- Monitoring and Reliability: Build robust monitoring and alerting for model health and latency, including service-level metrics for the data retrieval and inference gateway layers. Define SLOs and own on-call and incident response for the serving layer.
- Model Quality and Evaluation: Build automated evaluation and quality monitoring into the deployment path, regression and drift detection, offline/online evaluation, and LLM output quality checks, so models stay healthy long after launch.
- Research-to-Production Enablement: Partner with ML engineers and researchers to turn experimental models into production-ready services, providing self-service workflows and abstractions that let teams deploy safely and quickly without deep infrastructure expertise.
Unique Challenges
- Build and operate the production ML serving and Data Orchestration platform behind Mercari Group's AI and LLM features, serving tens of millions of users.
- Drive the strategy for model inference at scale, bridging the gap between complex data retrieval and fast-moving ML model inference to ensure high-performance, cost-effective service delivery.
- Shape Mercari's next-generation LLM serving stack, from inference optimization (quantization, dynamic batching, KV caching) to the evaluation and execution infrastructure needed for emerging agentic AI workloads.
Qualifications
- Required Experience/Skills
- Shared belief in the mission and values of Mercari Group.
- Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- 5+ years of software engineering experience, including proven experience in production MLOps: end-to-end model deployment, serving, and CI/CD in cloud environments.
- Experience designing and operating large-scale, high-availability distributed systems, including observability, SLO definition, and incident response.
- Strong experience in cloud-native infrastructure (Kubernetes, Docker).
- Proficiency in Python and infrastructure-as-code (Terraform).
- Excellent written and verbal communication.
- Preferred Experience/Skills
- Experience integrating ML serving with large-scale distributed data layers (e.g., data warehouses, wide-column stores, in-memory caches).
- Expertise in model inference optimization (TensorRT-LLM, quantization, JAX).
- Experience operating large-scale model inference gateways and orchestrators.
- 2+ years of hands-on experience operating GenAI/LLM workloads in production (e.g., LLM serving frameworks, token throughput and cost optimization).
- Experience building LLM evaluation, guardrail, or quality-monitoring pipelines (e.g., LLM-as-judge, golden datasets, drift detection).
- Experience with serving infrastructure for RAG or agentic AI workloads (vector search, tool-calling execution environments).
- Experience partnering closely with research or data science teams to bring research innovations into production.
- Master's or Ph.D. in a related technical field.
- Language
- English: Proficient (CEFR - B2)
- Japanese: Independent (CEFR - B2) optional
For details about CEFR, see here.
Learn More About Mercari Group
- Careers site:
- Mercan:
- Social media: X / Linkedin
Recruiting at Mercari
At Mercari Group, we value empathizing with and embodying the mission and values of the Group and each company. To promote the creation of an organization that maximizes the total amount of value exhibited by all members, we would like to understand the experience and skills of each candidate as accurately as possible.
Recruiting cycle at Mercari Group
- Application screening
- Skill assessment: For engineering positions, you will be asked to complete a skill assessment on HackerRank or GitHub. For non-engineering positions, you may be asked to complete an assessment depending on the position. (The timing of the assessment may coincide with the interview process.)
- Interview: The number of interviews may vary depending on the position.
- Reference check: We will ask for online references around the timing of the final interview.
- Offer: Offers will be determined carefully in consideration of the final interview and the reference check.
Learn more about our recruiting process here.
Equal Opportunity Hiring
Here at Mercari, we work to realize a world in which no one’s potential is limited by their background and everyone has the opportunity to freely create value. We also firmly believe that a mindset of Inclusion & Diversity is essential for us to achieve our mission.
This, of course, extends to our hiring practices as well. Mercari is committed to eliminating discrimination based on age, gender, sexual orientation, race, religion, physical disability, and other such factors so that anyone who shares our mission and values can join us, regardless of their background. For more details, please read our I&D statement.
Please read and acknowledge our Privacy Policy prior to submitting your application.
Skills
As published by workable
First name, Last name, Email, Phone, Address, Resume, Other, Please indicate your gender (We process your gender for statistical reasons) / 性別をご選択ください。 (メルカリでは上記の情報をダイバーシティー、インクルーシビティの促進のための統計目的で調査しております。), Please read and agree to the Recruitment Privacy Policy of the Mercari Group by checking "YES". (https://careers.mercari.com/en/privacy/) / メルカリグループの採用活動におけるプライバシーポリシーをご確認の上、同意(YES)をお願いします。(https://careers.mercari.com/jp/privacy/), Do you have an employment history with Mercari Group? / 過去にメルカリグループでの勤務経験がありますか?, Have you ever directly or indirectly applyed to any positions of Mercari Group? / 過去にご自身またはエージェントを通じてメルカリ グループに応募されたことはありますか?, How did you hear about this position? / この募集をどのように知りましたか?, If your application is a referral from an employee of Mercari Group, please the employee's name. If your application is not a referral, please enter NA. / 今回の応募がメルカリグループ社員による紹介の場合、社員名を記入下さい。該当しない場合はNAとご記入下さい。, Please select your current status of residence in Japan from the list below. We will confirm if you need visa support. / 現在お持ちの日本における在留資格を下記からお選びください。ビザサポートが必要になるか確認をさせていただきます。, Mercari's values serve as the foundation of our culture and guide all of our decisions. We encourage you to review them before submitting your application. (https://careers.mercari.com/culture/) メルカリのバリューは私たちの文化の柱であり、行動指針となります。あらかじめ応募前にご確認ください。 (https://careers.mercari.com/culture/), Select the language(s) in which you would feel comfortable interviewing? / 面接時の希望言語をお知らせください。, Please select the appropriate level about your English language ability from below. Please check each criteria from our web site. https://careers.mercari.com/language/ 英語スキルについて該当するものをご選択ください。レベルの基準は採用サイトよりご確認ください。 https://careers.mercari.com/jp/language/, Please select the appropriate level about your Japanese language ability from below. Please check each criteria from our web site. (https://careers.mercari.com/language/) 日本語スキルについて該当するものをご選択ください。レベルの基準は採用サイトよりご確認ください。 (https://careers.mercari.com/jp/language/), In consideration with ensuring an appropriate assignment and wellness after joining the company, please let us know if you have any medical history related to mental health. 入社後の適正配置及び健康確保の観点からの質問です。メンタルの不調に関連した既往歴がある場合はお知らせください。 YES=ある、NO=ない, If you have chosen "YES" in the previous question, please share your current status. Please choose NA if you have answered "NO" in the previous question. 上記設問で「ある」を選択された場合、通院状況についてお知らせください。上記で「ない」を選択された場合は該当なしを選択ください。, To ensure your well-being and that any necessary accommodations are prepared after joining the company, please let us know if you have any disabilities. 入社後の適正配置及び健康確保の観点より、障がいをお持ちの場合はお知らせください。 YES=ある、NO=ない, Please feel free to share if there is anything the company should know to help ensure you can perform at your best. そのほか、良いパフォーマンスをする上で会社が知っておくべきことがある場合はお知らせください。, Please enter your websites, social media accounts, GitHub, etc. If you do not have any, please enter N/A. お持ちのウェブサイト、ソーシャルメディア、GitHubアカウント等があればご記入ください。お持ちでない場合は「なし」とご記入ください。, This position requires relocation to Japan. Is this something you are open to? Candidates already located in Japan please select "yes". このポジションは日本への転居が必要です。日本への転居は可能ですか。すでに日本在住の方は "YES"を選択してください。