freehire launches on Product Hunt on 26 August.

Follow →

Senior DevOps Engineer / SRE

Position: Senior Observability Engineer / DevOps/ SRE

Location: Thang Long I IP, Hanoi (Daily Bus from Hanoi Center)
Who we are

MOLEX, a brand under KOCH Industries with its headquarters in Lisle, Illinois (USA), is one of the world’s leading interconnector manufacturers. Molex creates connections for life by providing technologies that transform the future and improve lives.Since 2007, Molex Vietnam has operated its first factory in Thang Long I IP. We are now expanding with a new high tech facility in Thang Long II IP, Hung Yen, backed by a $145 million investment and spanning 100,000 m². This site will manufacture advanced products including high-speed connectors, cables, and optical devices.


Job Purpose

We are hiring a Senior Observability Engineer / Platform Owner to drive enterprise observability enablement and ongoing operational visibility across

critical digital manufacturing systems.

This role will lead incoming dashboard, log, trace, alerting, and telemetry requests from ETS, MES, UFE, APS, AIP, SAP, Teamcenter, and related teams, turning

them into practical, scalable, and maintainable observability solutions.

The role is expected to own requirement clarification, solution design, prioritization, quality review, and cross-team execution, while leading 2 interns on day-to-day delivery and maintenance.


Key Responsibilities

Serve as the single intake point for observability requests across ETS, MES, UFE, APS, AIP, SAP, Teamcenter, and related teams.

• Design and break down solutions covering dashboards, alerts, logs, traces, matrix configuration, and telemetry collection.

• Evaluate priority, delivery sequencing, dependencies, and effort, and maintain the backlog.

• Review deliverables from 2 interns to ensure quality, consistency, reuse, and maintainability.

• Handle complex topics such as cross-system tracing, log ingestion boundaries, metric definitions, and alert threshold design.

• Drive key initiatives such as HVR sync alerting, MES logs to Loki, Alloy standardization, and Teamcenter / SAP observability visibility.

• Build and maintain runbooks, handover documentation, templates, and configuration standards.

• Work closely with K8s / SRE, DBA, Windows host admins, and application owners to land observability solutions.

• Lead known issue management, top error log analysis, and rule design for alert-to-known-issue closure.

• Define use cases for distributed trace support dashboards, trace detail workflows, and AI-assisted analysis support.

• Govern the boundary between Grafana-managed alerts and data source managed alerts, including notification paths, Alert manager configuration

ownership, and template strategy.

• Design solutions for complex dashboard data models such as QMS multilabel queries, panel reuse, hidden variables, and GUI embedding constraints.

• Define onboarding and integration approaches for external metrics such as

Confluent Kafka metrics, Admin Tools-based metrics, HVR sync metrics, and IoT / Ignition health metrics.

• As a longer-term evolution target, help transition the observability platform toward AIOps capability by preparing usable data foundations for anomaly detection, AI-assisted RCA, and intelligent alerting.

• Drive the capture and organization of key AIOps inputs such as metrics, logs, alerts, application dependency trees / graphs, application change events, historical issues, and RCA records.

• Work with SRE, ETS, and application teams to establish feedback loops for auto-RCA outputs, operational knowledge capture, and future optimization.

• Identify which alerting, log analysis, and known issue scenarios are suitable for future anomaly detection, intelligent analysis, or agent-assisted troubleshooting, without disrupting current observability foundation priorities.


Qualifications

Strong hands-on experience with Grafana, Prometheus, Loki, Tempo, or similar observability platforms.

• Solid experience in dashboard design, PromptQL or equivalent query languages, metric modeling, and alert design.

• Good understanding of how logs, traces, and metrics work together in end-to-end troubleshooting.

• Experience in requirement clarification, solution design, and technical coordination across teams.

• Ability to read and adjust telemetry and collection configuration such as Alloy, exporters, scrape jobs, and remote write.

• Strong documentation, review, and mentoring skills.

• Understanding of Open Telemetry, Open Observe, collectors / gateways, distributed trace querying, and log correlation analysis.

• Familiarity with Grafana contact points, webhooks, Alertmanager Config, Teams / Power Automate style notification workflows.

• Ability to handle complex dashboard data patterns, including SQL joins, panel reuse, hidden variables, iframe constraints, and multi-environment

data source design.

• Familiarity with telegraf, custom SQL metrics, ScrapeConfig, and third-party metrics onboarding approaches.

• Understanding of AIOps prerequisites such as data quality, historical RCA capture, event correlation, change event ingestion, and usable operational knowledge bases.

• Basic understanding of anomaly detection, intelligent alerting, AI-assisted RCA, and agentic workflow concepts, with the ability to judge when they are appropriate and when they are premature.

Nice to have:

• Experience in manufacturing or enterprise platforms such as MES, APS,SAP, Teamcenter, AIP, or UFE.

• Experience with Loki migrations, OpenTelemetry, distributed tracing, or alert governance.

• Experience leading interns or junior engineers.

• Experience with runbooks, KT materials, known issue repositories, or RCA

knowledge capture.

• Experience with Grafana webhooks, GenAI webhooks, known issue automation, or support dashboards.

• Experience with QMS / reporting-style dashboards, SQL optimization, or GUI integration.

• Experience onboarding Admin Tools, infra-monitor-tools, HVR sync, or

Confluent Kafka metrics.

• Experience with AIOps, anomaly detection, auto-RCA, LLM / GenAI assisted operations, event correlation, or operational knowledge management.

• Experience organizing application dependency graphs, change events, incident datasets, or RCA datasets for intelligent operations use cases.


Soft Skills & Personal Attributes

• Strong analytical and logical thinking for solving complex observability and operational problems.

• Excellent collaboration and communication skills for working across business, platform, application, DBA, and SRE stakeholders.

• Strong ownership mindset and the ability to drive progress in ambiguous, dependency-heavy situations.

• Strong prioritization, time management, and task decomposition skills.

• Ability to mentor interns and maintain review quality without becoming a delivery bottleneck.

• Passion for standardization, maintainability, continuous improvement, and high-quality operational execution.


🌟 Why Join Us?

· Modern American-standard work environment

· PBM culture that fosters collaboration and personal growth.

· Full salary during probation

· 24/7 occupational safety insurance from day one

· Premium health insurance after just 2 months

· Most Saturdays off

· Accommodation

· Global career development via KOCH’s Talent Exchange Program

· Annual Summer Trip

· Good-quality meals daily

· Total compensation is offered based on total contribution (including base salary, allowances, overtime, 13th-month salary, performance bonus, and spot bonus, service award).


At Koch companies, we are entrepreneurs. This means we openly challenge the status quo, find new ways to create value and get rewarded for our individual contributions. Any compensation range provided for a role is an estimate determined by available market data. The actual amount may be higher or lower than the range provided considering each candidate's knowledge, skills, abilities, and geographic location. If you have questions, please speak to your recruiter about the flexibility and detail of our compensation philosophy.


See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available