Aufbau und Weiterentwicklung der AI-Plattform auf Kubernetes
Deployment und Betrieb von Large Language Models (LLMs) auf GPU-Servern
Aufbau und Betrieb performanter Model-Serving- und Inference-Services
Etablierung von MLOps-Standards sowie Betrieb und Weiterentwicklung der on-prem MLOps-Plattform (u. a. ClearML, OpenShift)
Aufbau von LLMOps-Workflows für produktive GenAI-Services, einschließlich Serving-/Runtime-Konfiguration, Guardrails & Policies, Evaluierung und Regressionstests sowie Kosten- und Qualitäts-Monitoring
Containerisierung und Deployment von AI-Services mit Docker und Kubernetes
Skalierung und Optimierung von GPU- und CPU-Workloads
Entwicklung stabiler AI-APIs und Services für interne Anwendungen
Aufbau von Monitoring, Logging und Observability für AI-Systeme
Automatisierung von Infrastruktur-, Plattform- und Model-Releases (CI/CD,...