I work where GenAI architecture, product constraints and hands-on engineering meet.

I'm Daniele Moltisanti, an AI engineer and strategy lead based in Milan, Italy. I work on GenAI systems broadly, with a particular focus on Local and on-device AI - and on deciding what should run Local, Hybrid or Cloud.

Principal AI Engineer & AI Strategy Lead · Local AI & On-device LLMsSky Italia
Daniele Moltisanti portrait

What I work on

Three areas, one recurring question: what architecture makes this AI workload useful, controllable and sustainable?

GenAI systems

LLM evaluation, retrieval, integration and production architecture - starting from the product problem rather than a model catalogue.

Local & on-device AI

Understanding what can move from cloud APIs to infrastructure and devices we control, and where that trade-off actually creates value.

AI engineering & strategy

Connecting technical choices with privacy, performance, operating constraints and the decisions teams need to make.

Why my perspective may be useful

The advisory point of view comes from building systems, operating inside enterprise constraints and documenting the trade-offs in public.

Enterprise AI delivery

I lead and design AI initiatives in complex production environments at Sky Italia.

Hands-on systems

I build the infrastructure and proving grounds myself: local inference, Android runtimes, speech, product applications and measurement tooling.

Public technical work

I use stAI tuned, GitHub and LinkedIn to document architecture trade-offs, experiments and what survives contact with real systems.

What I'm exploring now

Current technical questions I’m using to push the Local / Hybrid / Cloud boundary.

Unified Memory Latency & Quantization Limits

Benchmarking 70B parameter models on M-series unified memory to determine the exact boundary where quantized open-weight inference matches frontier API responsiveness under sustained load.

Active Benchmarking

Android NPU Acceleration & Cross-App IPC

Prototyping system-level AIDL inference services on Android to enable lightweight local models to serve multiple applications without redundant RAM consumption.

Prototype Stage

Model Context Protocol (MCP) in Local Environments

Designing composable agent workflows where local models interact securely with private file systems and developer databases via standardized protocols without data leakage.

Architectural Design

Selected track record

Enough context to understand where the perspective comes from. LinkedIn has the full chronology.

Sky Italia2022 - Present

Data Scientist Manager / Principal AI Engineer & AI Lead Strategy

Directing enterprise AI and data science initiatives across subscriber platforms, content intelligence, and generative AI platforms serving millions of active users.

Independent Local AI Stack2024 - Present

AI Systems Architect & Open-Weight Researcher

Architected and built a comprehensive open-weight Local AI ecosystem spanning reusable inference servers, mobile harnesses, reference applications, and network observability.

stAI tuned2023 - Present

Founder & Managing Director

Founded an independent technical platform delivering deep architectural breakdowns, practical AI roadmaps, and no-hype guides for technical leaders and builders.

Have an AI workload to figure out?

Tell me what you're trying to run and what constraints matter. I'll start from the workload, not the model.