We are building an engineering AGI. We founded P-1 AI with the conviction that the greatest impact of artificial intelligence will be on the built world—helping mankind conquer nature and bend it to our will. Our first product is Archie, an AI engineer capable of quantitative and spatial reasoning over physical product domains that performs at the level of an entry‑level design engineer. We aim to put an Archie on every engineering team at every industrial company on earth.
We're looking for an experienced engineer to take ownership of LLM training operations across our applied research team. Your focus will be on making large‑scale GPU training run reliably, efficiently, and fast on a dedicated mid‑size GPU cluster and possibly on cloud platforms as well. You'll work closely with researchers and ML engineers developing new models and agentic systems, ensuring their experiments scale smoothly across multi‑node GPU clusters. From debugging NCCL deadlocks to optimizing FSDP configs, you'll be the go‑to person for training infrastructure and performance.
Mid‑Senior level
Full‑time
Engineering and Information Technology
Software Development