锦鲤求职

月之暗面

Research Scientist / Engineer – Agentic RL

锦鲤会替你打开官网网申、按简历自动填表提交,你只需要在关键步骤确认。

岗位描述

Research Scientist / Engineer – Agentic RL

Locations: Beijing · Shanghai · Shenzhen · Singapore · Silicon Valley

Mission

Invent the next leaps in agentic, multi-agent RL—algorithms, environments, infrastructure—that will carry Kimi K2.7 and its successors into real-world autonomy.

Core Requirements

• Deep mastery of RL: policy-gradient, actor-critic, self-play, meta-RL, MARL.

• Proven delivery of large-scale training & serving stacks (thousands of GPUs, low-latency rollouts).

• Expert-level systems coding (Python/C++/Rust) and distributed infra design.

• Track record of shipping or publishing on agentic RL, tool-use agents, or multi-agent systems.

Nice-to-have

• Experience with code-execution sandboxes, program synthesis, or simulator security.

Apply with GitHub, papers, or a concise write-up of your coolest RL system.

Welcome partners with SOTA capabilities in one of the above aspects to join us.

研究科学家 / 工程师 – Agentic RL

Moonshot AI · Kimi K2.7 及下一代模型

工作地点:北京 · 上海 · 深圳 · 新加坡 · 硅谷

使命

在 Agentic、多智能体 RL 的算法、环境与基础设施上取得突破,把 Kimi K2.7 及后续模型的自主能力推向真实世界。

核心要求

• 精通 RL:策略梯度、Actor-Critic、Self-Play、Meta-RL、MARL。

• 具备千卡级训练与低延迟推理的落地经验。

• 系统级编程与分布式架构设计能力(Python / C++ / Rust)。

• 有 Agentic RL、工具调用或多智能体系统的发表或上线记录。

加分项

• 代码沙箱、程序合成、仿真器安全相关经验。

请附上 GitHub、论文或你最自豪的 RL 项目简介。

欢迎在以上某一方面具备SOTA能力的小伙伴加入

人工智能方向的申请准备

先核对岗位要求与自己的经历,再准备档案、投递和面试。

阅读求职步骤与示例 →