Production Engineer
【關於這個角色】
Raccoon AI 是新興的生成式 AI 客服解決方案提供商,專注協助電子商務與網路平台企業,自動化處理客服訊息。
我們的產品運行在正式 SaaS production environment,服務企業客戶並承諾 SLA。隨著客戶與系統規模持續成長,我們希望建立更成熟的 Production Engineering 機制,讓 production issue 能被快速定位、處理與預防,同時降低產品開發團隊頻繁被 incident 打斷的情況。
我們正在尋找一位 Production Engineer,負責 production issue 的第一線工程判斷、debugging 與修復,並持續改善整體服務穩定性。
【工作內容】
負責 SaaS production issue 的 engineering triage、debugging 與問題定位
根據 SLA 與 severity 判斷事件優先級,協助快速恢復服務
查看 application log、database、API request、queue、monitoring 等資訊,找出 root cause
Reproduce 客戶回報的問題,判斷是產品 bug、資料問題、第三方服務異常或 infrastructure issue
能直接處理的 application bug,完成 hotfix、測試與 release
當問題需要深入 domain knowledge 時,整理完整技術資訊後 escalation 給對應 Feature Engineer
協助處理 Rails application、Python AI backend、API integration 與第三方服務相關問題
建立與維護 incident runbook、debugging tools、monitoring 與 alerting
分析 recurring incidents,找出系統性問題並推動改善
與 CS、PM、RD 協作,建立清楚的 incident response 與 escalation process
參與 postmortem,降低相同 production issue 再次發生的機率
【你需要具備】
3 年以上 Backend / Full-stack / Production Engineering 相關經驗
熟悉至少一種 Backend 技術棧,例如 Ruby on Rails、Python、Node.js 等
具備良好的 debugging 能力,能快速閱讀陌生 codebase 並定位問題
熟悉 SQL,能獨立進行 production database issue investigation
熟悉 REST API、Webhook、第三方 API integration 等常見 SaaS 架構
熟悉 log、monitoring、error tracking 等 production debugging 工具
理解 HTTP、network、database、queue、cache 等基本系統概念
能在資訊不完整的情況下快速整理問題、建立 hypothesis 並逐步排除
對 production reliability、incident handling 與 root cause analysis 有高度興趣
能與不同角色合作,清楚說明問題範圍、影響程度與處理進度
【加分條件】
Ruby on Rails 或 Python production experience
AWS / GCP / Kubernetes / Docker 經驗
熟悉 PostgreSQL、Redis、message queue
使用過 Datadog、Grafana、Sentry、CloudWatch 或類似 observability 工具
有 SaaS、B2B enterprise service 或 SLA environment 經驗
有 on-call、incident response、postmortem 經驗
有 AI / LLM application production experience
曾處理高流量 API、distributed system 或 third-party integration issue
【這個角色和一般 RD 有什麼不同】
一般 Product Engineer 的主要任務是開發新功能與產品 roadmap。Production Engineer 的主要任務則是確保已經上線的服務穩定運作,包括:
快速處理 production incident
縮短 MTTR
降低 SLA violation
找出 recurring issue 並消除 root cause
減少 Feature Engineer 被臨時 production issue 打斷
這不是單純「修 Bug」的角色,而是對 production system 有高度 ownership 的工程職位。
【我們重視的特質】
我們希望你看到 production issue 時,第一個反應不是追究這是誰寫的,而是思考「問題在哪裡、影響多少人、怎麼最快恢復、以及怎麼讓它不要再發生。」
如果你喜歡 troubleshooting、追 log、查 API、分析 database、理解整個系統如何運作,甚至覺得找出 tricky production bug 很有成就感,這個角色會非常適合你。
Skills
As published by ashby · 5 questions
Basics
Name, Email, Resume
Short answers (3)
- English Name/Prefer Name
- Contact Number
- Do you know any current employees at the company?
Pick from a list (2)
- Are you legally authorized to work in Taiwan?
- Consent to the Collection, Processing, and Use of Personal Data optional