岗位描述
Partner with cross-functional teams to define, build and roll out a scalable Generative AI agent platform for priority banking use cases from pilot to production. Own the end-to-end platform engineering: design and implement the agent runtime, orchestration, tool/function calling, prompt and workflow management, memory/state handling, and integration patterns—leveraging frameworks such as LangChain/LangGraph (where appropriate) while avoiding unnecessary vendor/framework lock-in. Establish platform reliability and quality by building in evaluation and testing (offline/online), guardrails, observability (logs/traces/metrics), and clear SLOs/SLAs to ensure accuracy, Continuously monitor and optimise platform performance in production, including model selection, prompt/agent tuning, retrieval quality, caching, throughput, and resilience; drive root-cause analysis and preventative improvements. Embed security, privacy and compliance by design: ensure controls for data handling, access management, encryption, auditability, model risk considerations, and safe tool use; align to internal governance and relevant regulatory expectations. Provide technical leadership and hands-on support: troubleshoot complex issues, support onboarding of delivery teams, create runbooks, and act as an escalation point for incidents and production support. Maintain strong platform documentation and enablement: publish reference architectures, APIs/SDKs, templates, best practices, and developer guidance to accelerate adoption and consistent delivery across teams. Track and evaluate emerging GenAI/agent capabilities (e.g., new orchestration patterns, evaluation methods, safety techniques, model/tooling advances) and translate them into pragmatic platform enhancements with clear value, risk and cost trade-offs. Education & English: Bachelor's degree or above in Computer Science/Engineering (or related); Good verbal communication in English. Software Engineering: 3+ years' hand-on coding practice, strong Python or Java or Golang, with usage of at least one major cloud (AWS/Azure/GCP); ability to deploy and operate services in production (Docker/Kubernetes preferred). GenAI / LLM Delivery: 1+ years building and shipping LLM/GenAI applications; knowledge with agentic workflows and tool/function calling (multi-agent systems a plus). SME in at least one area (Performance Monitoring, Agent Governance Control, Agent Identity, Harness, Agent SDK, Agent Policy / Guardrail) Be able to demonstrate a strong vision for the future of agent platforms. Be able to articulate concrete best practices for creating skills, agents, or MCP tooling. Problem-solving & commitment, strong ability to develop creative solutions to complex challenges, working effectively with technical and non-technical stakeholders. Demonstrates an AI-native mindset by applying AI-driven approaches, including coding assistants, to improve productivity, quality, and engineering best practices. RAG/MCP preferred: Practical RAG experience (retrieval, chunking, embeddings, vector databases such as Milvus/FAISS/Elastic/Pinecone)