Overview:
Ciklum is seeking an Artificial Intelligence / Machine Learning Engineer to join our team on a full-time basis in Ukraine (Kyiv).
About the project:
You will contribute to the Legacy Code Semantic Documentation Project, a 26-week enterprise initiative for a global industrial automation leader. The goal is to engineer an automated, AI-assisted pipeline to generate structured, system-level Markdown documentation directly from an undocumented codebase (~400K lines of code) spanning IEC 61131-3 languages and ANSI C/C++.
Technical environment:
Work within a dedicated, secure tenant using Tree-sitter for AST parsing, Neo4j knowledge graphs, Qdrant vector search, and self-hosted open-weight LLMs (Llama 3.1 / Mixtral). Outputs align with EU Cyber Resilience Act requirements for SBOM and source-level traceability.
Responsibilities
Develop and maintain multi-pass AI pipelines with self-hosted models
Integrate RAG retrieval using Qdrant and graph queries with Neo4j into unified prompt contexts
Support evaluation pipelines with automated NLI checks to verify factual accuracy
Design prompts and templates to generate structured Markdown assets
Collaborate with Data Engineers and QA to ensure outputs meet code baselines
Requirements
3+ years in commercial software development or data science, with 1.5+ years building and deploying practical AI/LLM applications
Overview:
Ciklum is seeking an Artificial Intelligence / Machine Learning Engineer to join our team on a full-time basis in Ukraine (Kyiv).
About the project:
You will contribute to the Legacy Code Semantic Documentation Project, a 26-week enterprise initiative for a global industrial automation leader. The goal is to engineer an automated, AI-assisted pipeline to generate structured, system-level Markdown documentation directly from an undocumented codebase (~400K lines of code) spanning IEC 61131-3 languages and ANSI C/C++.
Technical environment:
Work within a dedicated, secure tenant using Tree-sitter for AST parsing, Neo4j knowledge graphs, Qdrant vector search, and self-hosted open-weight LLMs (Llama 3.1 / Mixtral). Outputs align with EU Cyber Resilience Act requirements for SBOM and source-level traceability.
Responsibilities
Develop and maintain multi-pass AI pipelines with self-hosted models
Integrate RAG retrieval using Qdrant and graph queries with Neo4j into unified prompt contexts
Support evaluation pipelines with automated NLI checks to verify factual accuracy
Design prompts and templates to generate structured Markdown assets
Collaborate with Data Engineers and QA to ensure outputs meet code baselines
Requirements
3+ years in commercial software development or data science, with 1.5+ years building and deploying practical AI/LLM applications