Technical Support Engineer (LatAm)
We are looking for Technical Support Engineers who combine traditional IT Helpdesk responsibilities with application and production troubleshooting skills.
The role does not require regular software development. However, candidates must be able to investigate software issues, analyze application logs, follow technical playbooks, resolve issues that do not require code changes, and escalate more complex problems to the engineering team with sufficient technical context.
- Strong troubleshooting and analytical skills.
- Practical experience with Datadog Logs or a comparable logging and monitoring platform.
- Ability to search, filter, and analyze large volumes of application logs.
- Ability to read and interpret error messages, application logs, and basic stack traces.
- Experience with user account provisioning, permissions, authentication, and access management.
- Experience investigating software issues before escalating them to developers.
- Understanding of incident triage, prioritization, and escalation procedures.
- Ability to follow production runbooks and technical playbooks accurately.
- Ability to distinguish between:
• user or access issues;
• configuration issues;
• environment or infrastructure issues;
• application defects requiring a code change.
- Basic understanding of APIs, webhooks, web applications, and multi-service application workflows.
- Strong written and spoken English.
- Ability to work independently during the required US Eastern Time schedule.
- Careful and responsible approach to production systems and sensitive user access.
Preferred Skills
- Experience supporting SaaS or multi-tenant software platforms.
- Familiarity with AWS-based environments.
- Experience troubleshooting SMS, voice, or push notification workflows.
- Familiarity with Twilio or similar communication providers.
- Basic understanding of Node.js services, MongoDB, and Angular applications.
- Experience working with application engineering, SRE, or DevOps teams.
- Experience supporting healthcare, telehealth, or another regulated software product.
- Experience creating or improving internal support documentation and runbooks.
Service-Level Expectations
First Response Time
Response within 15 minutes: at least 80%
Response within 60 minutes: at least 95%
Blended FRT performance: approximately 92%
Resolution
Median resolution time: less than 12 hours
First-contact resolution rate: at least 75%
Delays caused by customer responses, engineering dependencies, or other external factors should be documented and reviewed individually.
- Create and manage user accounts across internal systems and application environments.
- Provision user access for newly created environments.
- Support laptop, authentication, permissions, access, and general IT-related issues.
- Investigate application and production issues using Datadog logs and monitoring tools.
- Search, filter, and analyze large volumes of application logs.
- Correlate events using timestamps, user information, tenant or environment details, request IDs, and other available identifiers.
- Follow documented troubleshooting playbooks and production runbooks.
- Resolve operational, configuration, access, and environment issues that do not require a code change.
- Execute documented production recovery or remediation procedures when an approved playbook is available.
Escalate issues to the appropriate on-call engineer when:
- a code change is required;
- no applicable playbook exists;
- the issue cannot be resolved safely within the Support Engineer’s permissions or scope.
Provide complete escalation details, including:
- affected user, tenant, or environment;
- issue symptoms and error messages;
- relevant logs and timestamps;
- troubleshooting steps already performed;
- observed results;
- business or user impact.
- Track incidents and support requests through the established ticketing process.
- Communicate clearly with users, engineers, and other stakeholders throughout the investigation.
- Contribute to troubleshooting documentation, knowledge base articles, and support playbooks.
- 5 sick days per year;
- 3 paid company holidays per year;
- Access to therapist and psychologist support for mental well-being;
- Compensation for courses and conferences;
- English courses.
The role does not require regular software development. However, candidates must be able to investigate software issues, analyze application logs, follow technical playbooks, resolve issues that do not require code changes, and escalate more complex problems to the engineering team with sufficient technical context.
Requirements
- Previous experience in Technical Support, Application Support, Production Support, IT Support, or a similar role.- Strong troubleshooting and analytical skills.
- Practical experience with Datadog Logs or a comparable logging and monitoring platform.
- Ability to search, filter, and analyze large volumes of application logs.
- Ability to read and interpret error messages, application logs, and basic stack traces.
- Experience with user account provisioning, permissions, authentication, and access management.
- Experience investigating software issues before escalating them to developers.
- Understanding of incident triage, prioritization, and escalation procedures.
- Ability to follow production runbooks and technical playbooks accurately.
- Ability to distinguish between:
• user or access issues;
• configuration issues;
• environment or infrastructure issues;
• application defects requiring a code change.
- Basic understanding of APIs, webhooks, web applications, and multi-service application workflows.
- Strong written and spoken English.
- Ability to work independently during the required US Eastern Time schedule.
- Careful and responsible approach to production systems and sensitive user access.
Preferred Skills
- Experience supporting SaaS or multi-tenant software platforms.
- Familiarity with AWS-based environments.
- Experience troubleshooting SMS, voice, or push notification workflows.
- Familiarity with Twilio or similar communication providers.
- Basic understanding of Node.js services, MongoDB, and Angular applications.
- Experience working with application engineering, SRE, or DevOps teams.
- Experience supporting healthcare, telehealth, or another regulated software product.
- Experience creating or improving internal support documentation and runbooks.
Service-Level Expectations
First Response Time
Response within 15 minutes: at least 80%
Response within 60 minutes: at least 95%
Blended FRT performance: approximately 92%
Resolution
Median resolution time: less than 12 hours
First-contact resolution rate: at least 75%
Delays caused by customer responses, engineering dependencies, or other external factors should be documented and reviewed individually.
Responsibilities
- Provide technical support during US Eastern Time business hours.- Create and manage user accounts across internal systems and application environments.
- Provision user access for newly created environments.
- Support laptop, authentication, permissions, access, and general IT-related issues.
- Investigate application and production issues using Datadog logs and monitoring tools.
- Search, filter, and analyze large volumes of application logs.
- Correlate events using timestamps, user information, tenant or environment details, request IDs, and other available identifiers.
- Follow documented troubleshooting playbooks and production runbooks.
- Resolve operational, configuration, access, and environment issues that do not require a code change.
- Execute documented production recovery or remediation procedures when an approved playbook is available.
Escalate issues to the appropriate on-call engineer when:
- a code change is required;
- no applicable playbook exists;
- the issue cannot be resolved safely within the Support Engineer’s permissions or scope.
Provide complete escalation details, including:
- affected user, tenant, or environment;
- issue symptoms and error messages;
- relevant logs and timestamps;
- troubleshooting steps already performed;
- observed results;
- business or user impact.
- Track incidents and support requests through the established ticketing process.
- Communicate clearly with users, engineers, and other stakeholders throughout the investigation.
- Contribute to troubleshooting documentation, knowledge base articles, and support playbooks.
Working conditions
- 12 vacation days per year;- 5 sick days per year;
- 3 paid company holidays per year;
- Access to therapist and psychologist support for mental well-being;
- Compensation for courses and conferences;
- English courses.