岗位描述
工作职责:
1. Build the operations team, processes, and service stability goals.
搭建运维团队、运维流程与服务稳定性目标。
2. Own cloud infrastructure, Kubernetes, CI/CD, and observability systems.
负责云基础设施、Kubernetes、CI/CD 与可观测性体系建设。
3. Establish monitoring, incident management, and capacity planning.
建立监控告警、事件管理与容量规划机制。
4. Implement security hardening, backup/recovery, and compliance controls.
落实安全加固、备份恢复与合规要求。
5. Manage and optimize cloud costs.
管理并持续优化云成本。
6. Collaborate with Engineering, AI, and Data, and report metrics and risks to the CTO.
与 Engineering、AI、Data 协作,并向 CTO 汇报运维指标与风险。
任职资格:
1. 8+ years of IT operations, SRE, or DevOps experience, including 3+ years in leadership.
8 年以上 IT 运维、SRE 或 DevOps 经验,其中 3 年以上团队管理经验。
2. Experience with major cloud platforms, Kubernetes, and container technologies.
熟悉主流云平台、Kubernetes 与容器化技术。
3. Proficiency in CI/CD, monitoring, logging, and observability.
熟悉 CI/CD、监控告警、日志与可观测性体系。
4. Strong incident management, security, and compliance awareness.
具备较强的事件管理、安全与合规意识。
5. Proficiency in English and Chinese as working languages; English is required.
中英文均可作为工作语言,英语为必备项。
6. Strong cross-team collaboration and communication skills.
具备较强的跨团队协作与沟通能力。
7. Cloud/SaaS, fintech compliance, FinOps, or 0-to-1 team building experience is preferred.
具备云服务/SaaS、金融科技合规、FinOps 或从 0 到 1 建团队经验者优先。