EXPERTISE · TIMELINE

Our engineering track record
2003 → 2026

21 years of building critical systems — from microcontroller firmware to sovereign LLM infrastructure.

05 — ENGINEERING DNA

21 years of engineering.
Hardware → AI → Production.

From bare-metal PCB to sovereign LLM — every layer mastered in-house, nothing outsourced. One technical point of contact who understands the full stack.

FOUNDER — Ali Abdel Aziz
Technical Curriculum
21
years
0
outsourced
2003
First real-time industrial firmware
C · RTOS · microcontrollers · sensor interfaces
Expert
2007
PCB & FPGA hardware architecture — international export
Altium · Xilinx · MIPI · power electronics
Expert
2012
High-frequency computer vision pipelines
PyTorch · TensorRT · ONNX · CUDA · IP cameras
Expert
2016
Distributed systems & multi-tenant SaaS — Dubai & Europe
Go · PostgreSQL · Kafka · K8s · AWS · microservices
Expert
2020
Sovereign on-premise LLM & fine-tuning — Morocco
Llama · Mistral · vLLM · LoRA · RAG · pgvector
Active
2024
Multimodal autonomous agents & Agentic Orchestration
LangGraph · n8n · MCP · vision + LLM + action loop
2026
CAPABILITY STACK
exp. · level
Hardware Design
PCB · FPGA · MIPI
18 ans
Embedded Systems
C · RTOS · Bare-metal
18 ans
Computer Vision
PyTorch · TensorRT · ONNX
13 ans
Agentic AI / LLMOps
LangGraph · vLLM · LoRA
6 ans
RAG & Knowledge Bases
pgvector · Qdrant · Milvus
4 ans
Workflow Automation
n8n · Temporal · Prefect
5 ans
SaaS & API Backend
Go · Node · PostgreSQL
12 ans
Frontend & Edge UI
React 19 · Vite · Workers
10 ans
Cloud & Infrastructure
AWS · K8s · Terraform
14 ans
Data & Analytics
Spark · dbt · ClickHouse
11 ans
IoT & Real-time Systems
MQTT · Kafka · Flink
9 ans
Security & Compliance
Zero Trust · CNDP
15 ans
Sovereign by default
No dependency on foreign APIs for core inference. Models run on-prem or in-region. Data never leaves the deployment perimeter.
Zero cold-start tolerance
Real-time inference pipelines designed for sub-10ms p99 latency. Model quantization and custom CUDA kernels where needed.
No vendor lock-in
Architecture is cloud-agnostic at the inference layer. AWS is the primary deployment target — but the models and runtime are 4YA's own IP.

Want the full story?

Meet the team behind 4YA — full bio, projects, and how we work.

Meet the team Back to home

How to read this track record

This timeline is not a résumé. It describes the capabilities the firm accumulated, period by period, and what each enables today. A skill only enters this list if it has served in production, not merely been learned.

The order matters: we came to artificial intelligence from below — from hardware, real time and systems that are not allowed to fail. This explains our architectural choices, notably the preference for edge processing.

What each period built

Embedded and real-time systems

Microcontroller firmware, board design, memory and power constraints. This period enforced a discipline you do not learn in business software: you cannot restart a device remotely, so the code must be right the first time.

Computer vision and distributed systems

Image processing on constrained hardware, then architectures distributed across several sites. This is where our V32 approach comes from: analyse as close to the camera as possible rather than uploading everything, because bandwidth and latency are scarce resources.

SaaS, cloud and operating at scale

Designing and operating multi-tenant platforms for international markets. This period brought what embedded work does not teach: continuous deployment, monitoring, version management and long-term operating cost.

Founding 4YA and Moroccan anchoring

4YA was founded in 2020 in Marrakech. The structuring choice was sovereignty: building systems whose data stays in Morocco, at a time when most offerings ran through foreign services. That choice shaped the entire stack that followed.

Proprietary products and sovereign AI

SecurePOS and the V32 module went into production, followed by open-weights model deployment on controlled infrastructure and autonomous agents in real operation. This is when the firm moved from services to publishing: we operate our own products, with real users.

What this accumulation enables today

Few teams cover the same span: from circuit board to language-model orchestration. That span is not a showcase argument — it has a practical consequence. When a vision system must run on a low-power unit, or a till must keep taking payments with the network down, the answer does not come from the application layer.

It is also what lets us say no. A share of the requests we receive is solved without artificial intelligence, through a process fix or a simple integration. Placing the problem at the right layer saves more time than any model.