In this interview, we speak with Uche Buzugbe, a UK-based Data Engineer and Machine Learning Engineer whose work spans cloud architecture, enterprise AI systems, and open-source innovation in multilingual machine learning. He shares insights into his career journey, technical focus, and contributions to advancing artificial intelligence systems for real-world impact.
Q: Can you briefly introduce yourself and your background?
I am a Data Engineer and Machine Learning Engineer with experience in designing scalable data systems, cloud-based architectures, and end-to-end machine learning workflows. My background combines software engineering, data analytics, and artificial intelligence, with a strong focus on building production-ready systems that transform raw data into actionable insights.
I hold a Master’s degree in Data Science from the University of Wolverhampton and a Bachelor’s degree in Industrial Physics from the University of Benin. Over the years, I have developed expertise across Python, SQL, Apache Spark, Kafka, TensorFlow, PyTorch, and cloud platforms such as AWS and Google Cloud Platform.
Q: What kind of work are you currently involved in?
I currently work as a Data Scientist and Machine Learning Engineer at AI Analytics Intelligence in the United Kingdom. My role involves building end-to-end machine learning pipelines, from data ingestion and preprocessing through feature engineering and model deployment using FastAPI.
I also develop NLP solutions using Hugging Face Transformers for tasks such as semantic search, classification, and summarization. In addition, I engineer real-time data streaming systems using Kafka and PySpark to handle high-volume data processing and enable fast, data-driven decision-making.
A significant part of my work involves designing scalable cloud-based machine learning systems using AWS and Google Cloud Platform to ensure performance, reliability, and scalability in production environments.
Q: Can you describe your previous experience in data engineering and AI systems?
Previously, I worked as a Senior Data Engineer and Scientist at Cara in the UK, where I was responsible for designing scalable data pipelines using PySpark and Apache Airflow for complex healthcare datasets.
I also implemented ETL workflows using AWS S3, EMR, and Lambda, improving data processing efficiency and enabling better access to structured datasets for analytics teams. My work included integrating machine learning models using TensorFlow and PyTorch for predictive analytics in healthcare environments.
Earlier in my career at WakaPadi, I worked as a Data Analyst and Technical Writer, where I developed ETL pipelines, analyzed structured and unstructured datasets using Python and SQL, and created technical documentation to support both technical and non-technical stakeholders. I also contributed to CI/CD workflows and containerized deployments using Docker and Kubernetes.
Q: You have also contributed to open-source AI. Can you tell us about that?
Yes. One of my key open-source contributions is a project called NaijaEval, a machine learning evaluation framework designed specifically for African language AI systems. It started in April 2026
NaijaEval addresses a major gap in natural language processing: the lack of reliable evaluation tools for low-resource and multilingual African languages. Many existing evaluation systems are designed for high-resource languages and fail to properly measure performance in code-switching and dialect-rich environments.
NaijaEval introduces evaluation metrics that assess hallucination detection, terminology preservation, and code-switching robustness. It supports languages such as Yoruba, Igbo, Hausa, Nigerian Pidgin, Swahili, Zulu, and Amharic.
The project is open-source and available on GitHub and PyPI under the Apache 2.0 license, allowing researchers and developers globally to extend and apply it in their own machine learning workflows. My goal with this project is to contribute to more inclusive and reliable AI evaluation systems for underrepresented languages.
Q: What drives your work in AI and data engineering?
I am motivated by the challenge of building systems that are not only technically scalable but also practically useful in real-world environments. I enjoy working across the full machine learning lifecycle from data engineering and model development to deployment and monitoring.
I am also particularly interested in how AI can be made more inclusive and accessible, especially for languages and regions that are often underrepresented in mainstream machine learning research.
Q: What are your long-term goals?
My long-term goal is to continue building scalable AI systems while contributing to the advancement of inclusive machine learning infrastructure. I aim to develop solutions that bridge the gap between research and production systems, particularly in areas like multilingual AI, cloud-native machine learning, and real-time data systems.
I also plan to expand my open-source contributions and collaborate with researchers and engineers working on similar challenges in artificial intelligence.