가우스랩스 · 인프라·DevOps
- 산업 AI·제조 데이터 분석
- 2020년 설립 (7년차)
Senior Site Reliability Engineer (KR)
이 브라우저에 저장돼요 · 최근 확인
무슨 일을 하고, 무엇을 요구하나요?
인프라·DevOps · 필수 · 경력 5년 이상
- 주요 업무
- Platform reliability and operations: Own platform-layer reliability across both environments.
- Monitoring and Alerting: Build and maintain robust monitoring and alerting for the infrastructure and platform layer to proactively identify and resolve issues before they impact the platform.
- Incident Response: Own incident first-response for the platform layer and participate in the on-call rotation to minimize downtime and restore service quickly.
- Automation: Develop automation tools and scripts to streamline operations, reduce manual effort, and enable engineering teams to operate their own services safely.
- 자격 요건
- Bachelor's degree in computer science, engineering, or a related discipline
- 경력 조건5+ years of industry experience as a Site Reliability Engineer or in platform/infrastructure engineering
- Hands-on experience operating Kubernetes in production (EKS preferred): cluster lifecycle, scheduling, autoscaling, resource management
- Experience with cloud platforms (AWS preferred) and containerization technologies (Docker, Kubernetes)
- 우대 사항
- Knowledge of AI/ML infrastructure and workloads.
- Knowledge of database technologies (MongoDB, PostgreSQL)
- Experience operating software in customer-managed (on-prem or customer-cloud) environments
- Exposure to manufacturing, semiconductor, or enterprise B2B customer environments
- 근무·전형
- Application reivew - Phone interview - Virtual onsite interview - VP interview/Core Value interview - CEO interview
공고 원문의 문장을 그대로 옮겼어요. 근무 조건과 전형은 바뀔 수 있으니 지원 전에 원문을 확인하세요.
가우스랩스, 어떤 회사일까요?
- 산업 AI·제조 데이터 분석
- 2020년 설립 (7년차)
제조 데이터와 AI를 결합해 공정 예측·계측을 지원하는 산업 AI 기업입니다. 2020년 미국에서 설립됐으며 서울과 팰로앨토 거점을 안내합니다.
가우스랩스 더 알아보기포지션 소개 · 공고 전문 읽기
Gauss Labs is an industrial AI company on a mission to revolutionize manufacturing with AI, starting with the semiconductor sector. Panoptes is an AI-based virtual metrology solution deployed in high-volume manufacturing fabs, helping customers improve yield, reduce costs, and accelerate production. Our software runs in our customers' own managed environments, and we're seeking a Site Reliability Engineer to own the reliability of the infrastructure and platform that Panoptes runs on. You will keep the platform available, performant, and scalable; own monitoring, alerting, incident first-response, and the on-call rotation; and build the automation and observability that let engineering teams operate their services safely.
Responsibilities
- Platform reliability and operations: Own platform-layer reliability across both environments.
In our internal cloud environment: full ownership — cluster health, resource management (CPU/memory/OOM), scheduling, autoscaling, Kubernetes/EKS lifecycle. In the customer environment: operate directly at the application-namespace level and for the customer-controlled cluster/node layer, diagnose and clearly communicate what's needed, and operate the platform within their setup, decisions, and constraints.
- Monitoring and Alerting: Build and maintain robust monitoring and alerting for the infrastructure and platform layer to proactively identify and resolve issues before they impact the platform.
- Incident Response: Own incident first-response for the platform layer and participate in the on-call rotation to minimize downtime and restore service quickly.
- Automation: Develop automation tools and scripts to streamline operations, reduce manual effort, and enable engineering teams to operate their own services safely.
- Capacity Planning: Forecast resource needs, optimize resource utilization, and ensure the platform infrastructure can handle increasing workloads.
- Deployment infrastructure: Build and maintain CI/CD pipelines and deployment infrastructure for the platform.
- Continuous Improvement: Drive a culture of continuous improvement by identifying opportunities to enhance platform reliability, performance, and efficiency.
Basic Qualifications
- Bachelor's degree in computer science, engineering, or a related discipline
- 5+ years of industry experience as a Site Reliability Engineer or in platform/infrastructure engineering
- Hands-on experience operating Kubernetes in production (EKS preferred): cluster lifecycle, scheduling, autoscaling, resource management
- Experience with cloud platforms (AWS preferred) and containerization technologies (Docker, Kubernetes)
- Experience with observability and alerting tools (Prometheus, Grafana, ElasticSearch, Jaeger)
- Experience with scripting languages (Python, Bash)
- Working knowledge of GitHub, GitHub Actions, and CI/CD concepts
- Strong problem-solving and troubleshooting skills
- Working proficiency in English for internal documentation and technical coordination
Preferred Qualifications
- Knowledge of AI/ML infrastructure and workloads.
- Knowledge of database technologies (MongoDB, PostgreSQL)
- Experience operating software in customer-managed (on-prem or customer-cloud) environments
- Exposure to manufacturing, semiconductor, or enterprise B2B customer environments
[Interview process]
Application reivew - Phone interview - Virtual onsite interview - VP interview/Core Value interview - CEO interview