GenAI systems
LLM evaluation, retrieval, integration and production architecture - starting from the product problem rather than a model catalogue.
I'm Daniele Moltisanti, an AI engineer and strategy lead based in Milan, Italy. I work on GenAI systems broadly, with a particular focus on Local and on-device AI - and on deciding what should run Local, Hybrid or Cloud.

Three areas, one recurring question: what architecture makes this AI workload useful, controllable and sustainable?
LLM evaluation, retrieval, integration and production architecture - starting from the product problem rather than a model catalogue.
Understanding what can move from cloud APIs to infrastructure and devices we control, and where that trade-off actually creates value.
Connecting technical choices with privacy, performance, operating constraints and the decisions teams need to make.
The advisory point of view comes from building systems, operating inside enterprise constraints and documenting the trade-offs in public.
I lead and design AI initiatives in complex production environments at Sky Italia.
I build the infrastructure and proving grounds myself: local inference, Android runtimes, speech, product applications and measurement tooling.
I use stAI tuned, GitHub and LinkedIn to document architecture trade-offs, experiments and what survives contact with real systems.
Current technical questions I’m using to push the Local / Hybrid / Cloud boundary.
Benchmarking 70B parameter models on M-series unified memory to determine the exact boundary where quantized open-weight inference matches frontier API responsiveness under sustained load.
Prototyping system-level AIDL inference services on Android to enable lightweight local models to serve multiple applications without redundant RAM consumption.
Designing composable agent workflows where local models interact securely with private file systems and developer databases via standardized protocols without data leakage.
Enough context to understand where the perspective comes from. LinkedIn has the full chronology.
Directing enterprise AI and data science initiatives across subscriber platforms, content intelligence, and generative AI platforms serving millions of active users.
Architected and built a comprehensive open-weight Local AI ecosystem spanning reusable inference servers, mobile harnesses, reference applications, and network observability.
Founded an independent technical platform delivering deep architectural breakdowns, practical AI roadmaps, and no-hype guides for technical leaders and builders.
Tell me what you're trying to run and what constraints matter. I'll start from the workload, not the model.