Who You'll Work For
At Makersite, we're pioneering the future of sustainable product development and digital collaboration. As a leading platform for product lifecycle management (PLM), we empower companies to make smarter, more sustainable decisions across their entire supply chain. This role is a fixed, permanent position, and all successful applicants will receive a permanent employment contract regardless of location.
The Role
Develop and Deploy Production-Grade AI Systems
- Build, deploy, and maintain scalable APIs that serve AI/ML models in production environments
- Own end-to-end delivery of AI solutions, from prototyping to fully productionized systems
Design Advanced AI Architectures
- Design and implement Retrieval-Augmented Generation (RAG) systems using vector databases
- Build and orchestrate AI agents using frameworks such as LangGraph, CrewAI, or similar
- Evaluate and select appropriate large language models (LLMs) and foundation models based on specific use cases
Optimize Performance and Scalability
- Continuously optimize model inference for latency, cost efficiency, and throughput at scale
- Identify bottlenecks and implement improvements to ensure high-performing systems
Ensure Reliability and Observability
- Implement robust monitoring, logging, and alerting for deployed models and services
- Ensure system reliability, uptime, and performance through best practices in production ML systems
Establish Engineering Best Practices
- Build and maintain CI/CD pipelines for seamless testing, deployment, and iteration of AI systems
- Contribute to evolving engineering standards as the company transitions from rapid experimentation to more structured, scalable operations
Core Experience
- 5+ years of experience in Python, with ~10 years of overall software engineering or data experience (candidates transitioning from backend engineering are welcome)
- Proven experience building and deploying production-grade APIs (preferably with FastAPI)
- Hands-on experience working with large language models (LLMs) and LLMOps, including prompt engineering, fine-tuning, and evaluation (e.g., GPT, Claude)
- Strong experience fine-tuning open-source models (e.g., Hugging Face ecosystem)
- Practical experience designing and working with vector databases (e.g., Pinecone, Weaviate, Chroma, pgvector)
- Experience building AI agents using frameworks such as LangGraph, LangChain, CrewAI, or similar
- Solid understanding of model deployment and serving (e.g., vLLM, TGI, or managed endpoints)
- Experience with CI/CD pipelines and modern deployment practices (Docker, Kubernetes, GitHub Actions)
- Residing in and legally permitted to work in the EU
What We Offer
- Competitive Salary
- 30 Days Paid Time Off
- Remote-First Flexibility – Work from anywhere in the EU, with the option to collaborate in person at our offices in Stuttgart, Berlin (role dependent)
- Generous Learning & Development Budget
- Choose Your Ideal Work Equipment