Shield AI is seeking a Staff Data Platform Engineer to define and build the data foundation of the AI Factory. The Data Platform provides a unifying, knowledge-graph-centered API layer for human and agentic workflows.
Responsibilities:
- Develop a unifying Graph API: Lead the architecture and implementation of the knowledge graph and multi-modal API layer that serves as the backbone for human, service, and agentic workflows.
- Own DataOps infrastructure: Research, optimize, and maintain the storage, indexing, query, ingestion, and compute infrastructure used throughout the data lifecycle.
- Establish best-practices: Establish durable, best-practice patterns for schema modeling, relationships, lineage, and schema evolution.
- Turbocharge agentic data access: Build APIs that enable agents to retrieve structured, connected, and explainable context rather than relying only on keyword or vector similarity.
- Develop reference architectures: Establish recommended storage and compute profiles, deployment patterns, benchmarks, and operational guidance for both internal and customer-managed infrastructure.
- Advise downstream teams: Partner directly with autonomy, ML, test, infrastructure, product, and customer-facing teams to turn real workflows into reusable platform capabilities from modeling to integrations.
- Build first-party integrations: Deliver integrations that make important data easy to collect and aggregate, including data produced by simulations, test infrastructure, training systems, and edge devices.
- Improve developer experience: Create self-service APIs, SDKs, tools, examples, and diagnostics that make correct data modeling and ingestion the easiest path.
- Drive technical direction: Evaluate emerging data and AI infrastructure technologies, make principled build-versus-buy decisions, and guide implementation across team boundaries.
- Raise operational quality: Establish expectations for observability, performance, reliability, security, data integrity, disaster recovery, and lifecycle management.
Key outcomes:
- Human and agentic workflows use one coherent API for discovering data, traversing relationships, and accessing specialized payloads.
- Teams spend their time deciding how to model and use data rather than repeatedly deciding where and how to store it.
- Data produced at the edge, in simulation, during testing, and in training flows into reusable platform models with minimal integration friction.
- Portable and operational platform capabilities across all deployment environments.
- Downstream teams can adopt the platform through stable APIs and SDKs instead of custom point-to-point integrations.
Required qualifications:
- Significant experience designing and operating distributed data solutions, storage systems, or data-intensive backend services.
- Strong software engineering skills and a record of delivering production systems in languages such as Go and Python.
- Deep understanding of data modeling, API design, schema evolution, identity, consistency, indexing, query planning, and data lifecycle concerns.
- Experience working across multiple storage modalities, such as relational or graph databases, object storage, analytical or columnar systems, and file storage.
- Experience designing reliable ingestion and access paths for high-volume or operationally important data.
- Strong understanding of Kubernetes, Linux, networking, security, storage, observability, and distributed-systems fundamentals.
- Experience deploying data infrastructure across cloud or customer-managed environments using modern Infrastructure as Code and platform engineering practices.
- Ability to evaluate technologies through prototypes, benchmarks, operational requirements, and total lifecycle cost rather than feature lists alone.
- Experience defining architecture and technical standards while remaining hands-on in implementation and debugging.
- Demonstrated ability to collaborate with ML researchers, autonomy engineers, test teams, platform engineers, and product stakeholders.
- Clear technical communication and the ability to make complex data architecture understandable to both specialists and downstream users.

