岗位描述
Role Description
This position is part of the Digital Labs team within the Global Digital Technologies Department. The team brings together passionate professionals from research and development, artificial intelligence, data science, data engineering, software engineering, bioinformatics, and product management.
We uphold a spirit of craftsmanship to develop industry-leading digital platforms, data products, and AI solutions that accelerate drug discovery, process development, and biopharmaceutical manufacturing, ultimately improving the speed and quality of innovation in human health.
As a Data & AI Full-Stack Engineer, you will participate in the end-to-end lifecycle of data and AI solutions, covering data collection and engineering, analysis and modeling, algorithm development, application integration, deployment, and continuous improvement.
Key Responsibilities
1. Business Requirement Analysis and Solution Design: Understand business scenarios and requirements, clarify objectives and constraints, and translate real-world problems into feasible data and AI solutions.
2. Data Collection and Engineering: Collect, clean, integrate, and model structured and unstructured data from multiple sources. Develop reliable ETL/ELT workflows, data pipelines, analytical datasets, and reusable data assets.
3. Data Analysis, Visualization, and Communication: Conduct exploratory and statistical analysis, develop effective visualizations, and clearly communicate findings and recommendations to technical and non-technical stakeholders.
4. Algorithm and AI Development: Design, develop, and evaluate solutions based on business needs, including but not limited to:
o Machine learning models for classification, regression, time-series forecasting, and deep learning
o Generative AI applications, RAG pipelines, AI agents, and LLM workflows
o Mathematical optimization, scheduling, resource allocation, and heuristic algorithms
o Natural language processing, knowledge graphs, and other applied AI technologies
5. Solution Implementation and Productization: Convert data analysis and algorithm prototypes into reliable services, APIs, applications, or automated workflows. Support testing, deployment, monitoring, performance optimization, and issue resolution.
6. Software Engineering and Data Quality: Follow good engineering practices, including modular design, version control, testing, documentation, reproducibility, security, data quality, and responsible AI principles.
7. Technology Research and Continuous Improvement: Explore emerging data and AI technologies, evaluate their applicability, and apply suitable approaches to real business problems.
8. Cross-functional Collaboration: Collaborate with business teams, product managers, data scientists, data engineers, and software developers to deliver high-quality data products and AI solutions.
Role Requirements
1. Education: Master’s degree or above in Computer Science, Software Engineering, Data Science, Statistics, Mathematics, Operations Research, Bioinformatics, Computational Biology, or a related field. PhD candidates and outstanding bachelor’s degree candidates will also be considered.
2. Programming Foundation: Solid foundation in data structures, algorithms, and programming. Proficiency in Python and the ability to write clear, maintainable, and testable code.
3. Data Science and Machine Learning: Understanding of statistics, machine learning fundamentals, model evaluation, and experimental design. Familiarity with frameworks or libraries such as Scikit-learn, PyTorch, or TensorFlow.
4. Data Engineering: Proficiency in SQL and familiarity with relational databases. Understanding of data modeling, ETL/ELT, data quality, and data pipeline development. Experience with NoSQL databases, Spark, Flink, Kafka, Airflow, Dagster, or similar technologies is preferred.
5. AI and Algorithm Experience: Practical project, research, internship, or competition experience in one or more areas such as machine learning, deep learning, Generative AI, natural language processing, knowledge graphs, mathematical optimization, or operations research.
6. Engineering and Deployment: Familiarity with Git and basic software engineering practices. Experience with API development, Linux, Docker, cloud platforms, CI/CD, model deployment, or application development is preferred.
7. Problem-solving and Learning Ability: Strong analytical thinking, curiosity, and the ability to quickly learn unfamiliar business domains, technologies, and engineering practices.
8. Communication and Teamwork: Strong communication and teamwork skills, with the ability to collaborate effectively across functions and clearly explain technical ideas. Proficiency in Mandarin and English for professional communication is preferred.
9. Preferred Experience: Publications, patents, open-source contributions, technical competition achievements, or experience in delivering end-to-end data or AI projects are preferred. Experience or academic background in biopharmaceutical research, development, manufacturing, or life sciences is a plus.
岗位说明
本岗位隶属于全球数智科技部Digital Labs团队。团队汇聚了研发、人工智能、数据科学、数据工程、软件工程、生物信息及产品管理等领域的优秀人才。
我们秉承工匠精神,致力于打造行业领先的数字化平台、数据产品和人工智能解决方案,加速药物研发、工艺开发及生物制药生产,持续提升生命健康领域创新的速度与质量。
作为数据与算法全栈工程师,你将参与数据与人工智能解决方案的完整生命周期,覆盖数据采集与工程、数据分析与建模、算法开发、应用集成、部署及持续优化。
岗位职责
1. 业务需求分析与方案设计: 理解业务场景和项目需求,明确目标及约束,将实际业务问题转化为可落地的数据与人工智能解决方案。
2. 数据采集与数据工程: 对来自多种数据源的结构化及非结构化数据进行采集、清洗、整合和建模,开发可靠的ETL/ELT流程、数据管道、分析数据集及可复用的数据资产。
3. 数据分析、可视化与沟通: 开展探索性分析和统计分析,设计有效的数据可视化,并向技术及非技术相关方清晰传达分析结果和建议。
4. 算法与人工智能开发: 根据业务需求设计、开发和评估适合的算法解决方案,包括但不限于:
o 用于分类、回归、时间序列预测及深度学习的机器学习模型
o 生成式人工智能应用、RAG流程、AI智能体及大模型工作流
o 数学优化、排程、资源分配及启发式算法
o 自然语言处理、知识图谱及其他应用型人工智能技术
5. 方案实现与产品化: 将数据分析和算法原型转化为可靠的服务、API、应用或自动化工作流,并支持测试、部署、运行监控、性能优化及问题排查。
6. 软件工程与数据质量: 遵循良好的工程实践,包括模块化设计、版本管理、测试、文档、结果可复现性、安全、数据质量及负责任人工智能原则。
7. 技术研究与持续优化: 探索新的数据与人工智能技术,评估其适用性,并将合适的技术应用于实际业务问题。
8. 跨团队协作: 与业务团队、产品经理、数据科学家、数据工程师及软件开发人员紧密合作,交付高质量的数据产品和人工智能解决方案。
岗位要求
1. 学历背景: 计算机科学、软件工程、数据科学、统计学、数学、运筹学、生物信息、计算生物学或相关专业硕士及以上学历。博士及优秀本科毕业生亦可考虑。
2. 编程基础: 具备扎实的数据结构、算法和编程基础,熟练掌握Python,能够编写清晰、可维护且可测试的代码。
3. 数据科学与机器学习: 理解统计学、机器学习基础、模型评估及实验设计,熟悉Scikit-learn、PyTorch或TensorFlow等框架或工具。
4. 数据工程能力: 熟练掌握SQL并熟悉关系型数据库,理解数据建模、ETL/ELT、数据质量及数据管道开发。具备NoSQL数据库、Spark、Flink、Kafka、Airflow、Dagster或类似技术经验者优先。
5. 人工智能与算法经验: 在机器学习、深度学习、生成式人工智能、自然语言处理、知识图谱、数学优化或运筹学等一个或多个方向具备项目、科研、实习或竞赛经验。
6. 工程化与部署能力: 熟悉Git及基本软件工程实践。具备API开发、Linux、Docker、云平台、CI/CD、模型部署或应用开发经验者优先。
7. 问题解决与学习能力: 具备良好的逻辑分析能力、求知欲,以及快速学习陌生业务领域、新技术和工程实践的能力。
8. 沟通与团队合作: 具备良好的沟通和团队合作能力,能够开展跨团队协作并清晰表达技术思路。能够使用中文和英文进行工作沟通者优先。
9. 优先条件: 在相关领域拥有论文、专利、开源项目贡献、技术竞赛成果,或具备完整数据及人工智能项目交付经验者优先。具备生物医药研发、工艺开发、生产制造或生命科学相关经验或学术背景者优先。