Pure Imagination Studios is a diversified entertainment company focused on storytelling and immersive experiences. They are seeking an AI Engineer to identify and integrate generative AI tools into various workflows, enhancing content creation and supporting simulation delivery through collaboration with cross-functional teams.
Responsibilities:
- Develop and maintain the on-prem LLM stack: model selection, deployment, and the tradeoffs that get acceptable performance out of the available hardware (current spec includes RTX PRO 6000 Blackwell-class GPUs)
- Collaborate with cross-disciplinary teams to implement practical AI solutions for real-time experiences, production tools, and support systems
- Create and manage the retrieval system that connects the model to a custom domain corpus, and the strategies that keep the system honest when it doesn’t know
- Oversee the Stand up and tune the local speech pipeline (ASR in, TTS out) for the latency and naturalness a live simulation need
- Develop the integration surfaces between the AI and the rest of the product and support everything that talks to the AI to define what those conversations look like
- Support ethical and responsible use of AI technologies, balancing innovation with practical and repetitional risks
- Collaborate with cross-disciplinary teams to establish how accuracy and hallucination control are measured and accepted on this product
- Plan for concurrency: the system needs to serve many simultaneous participant interactions across the simulation
- Assistance with defining how knowledge stays current in an air-gapped environment with establishing a partnership with operations
- Partner with leadership to define and evolve AI-related strategy, standards, and governance across projects
- Performs other duties as assigned
Requirements:
- Associate's degree or equivalent from two-year college/technical school required
- Minimum 3 to 5 years of experience using automation, scripting, or machine learning tools to support creative, technical, or operational workflows
- Minimum 1 to 2 years of active hands-on work with modern generative AI tools (e.g., ChatGPT, Stable Diffusion, Whisper, ElevenLabs)
- Deployed and managed at least one production on-prem or air-gapped LLM system. Cloud-only experience is not sufficient
- Expertise with local LLM deployment: model selection, quantization, inference runtimes (llama.cpp, vLLM, TGI, or similar), and GPU resource management
- Proficiency in Building RAG pipelines against custom corpora in production, demonstrates a strong working knowledge of chunking, retrieval evaluation, and prompt design
- Successful track record of measuring hallucination on production systems, not just shipping systems that seem to work
- Working knowledge of local ASR and TTS toolchains and the latency tradeoffs in a speech pipeline
- Strong verbal and written communication skills, with the ability to interact effectively with internal teams, external vendors, and other stakeholders
- Ability to handle sensitive information with confidentiality and professionalism
- Bachelor's degree from a four-year college or university is preferred or equivalent combination of education and experience
- Relevant profession certifications preferred
- Expertise in Real-time interactive experiences, simulation systems, or game-engine-adjacent products
- Experience working with Multi-tenant LLM architectures where the same engine serves multiple products or personas
- Expertise in pragmatic about tool choice, selects the right tool for the job rather than the most interesting one