ABOUT · LOCAL AI

Daniele Moltisanti

I build Local AI systems to prove when AI can leave the cloud.

I design, build, and benchmark AI running on devices and infrastructure we control across desktop, mobile, and speech.

The goal: decide with evidence when Local, Hybrid or Cloud is the right architecture.

Principal AI Engineer & AI Strategy Lead · Local AI & On-device LLMsMilan, Italy
Daniele Moltisanti portrait
01 · BUILD

Reusable local infrastructure

Inference runtimes and system boundaries that applications can actually depend on.

02 · TEST

Real applications

Privacy, mobile and meeting workflows that expose constraints demos can hide.

03 · MEASURE

Performance & boundary evidence

Latency, memory, thermal and network signals that can change the architecture decision.

NOW · AUGUST 2026

What I’m testing now

Current work at the boundary between model capability and real execution constraints.

Apple Silicon MLX

Unified Memory Latency & Quantization Limits

Benchmarking 70B parameter models on M-series unified memory to determine the exact boundary where quantized open-weight inference matches frontier API responsiveness under sustained load.

Active Benchmarking
On-Device AI

Android NPU Acceleration & Cross-App IPC

Prototyping system-level AIDL inference services on Android to enable lightweight local models to serve multiple applications without redundant RAM consumption.

Prototype Stage
Agentic Architectures

Model Context Protocol (MCP) in Local Environments

Designing composable agent workflows where local models interact securely with private file systems and developer databases via standardized protocols without data leakage.

Architectural Design
Technical Writing

Zero-Hype Guide to Inference Economics & TCO

Authoring a comprehensive comparative guide on stAI tuned detailing when self-hosted and on-device inference becomes more cost-effective than frontier API calls.

Publishing Soon

HOW I DECIDE

No Local AI ideology. Just architectural evidence.

The same principles guide what I build, benchmark and recommend.

01STRATEGY

Local AI First ≠ Local AI Only

Every AI architecture choice is a trade-off among privacy, latency, runtime ownership, cost, and raw reasoning capacity. The objective is to find the boundary empirically: assigning tasks to Local, Hybrid, or Cloud based on verifiable workload requirements.

03PRIVACY

Architectural Data Boundaries

If sensitive data never leaves the user’s local memory boundary, the attack surface shrinks to near zero. By designing models as localized transformation primitives, we eliminate compliance friction and guarantee true data sovereignty.

04EVALUATION

Empirical Verification Over Leaderboards

A model scoring high on MMLU might suffer from terrible time-to-first-token, erratic JSON schema compliance, or extreme thermal throttling on mobile hardware. Systems must be tested against real domain inputs on actual deployment targets.

WHY I WORK THIS WAY

Research rigor, enterprise constraints, then back to the device.

The track record matters because Local AI is a systems problem, not only a model problem.

01
Politecnico di Milano · IIT
THE FOUNDATION

Rigorous Computer Engineering & Academic Research

Built a foundational respect for empirical metrics, systems performance, and computational boundaries.

02
Sky Italia · 2022 - Present
ENTERPRISE SCALE

Leading AI & Data Science in Production

Mastered the bridge between executive AI strategy, MLOps rigor, and large-scale subscriber impact.

03
stAI tuned · 2023 - Present
ZERO-HYPE DIVULGATION

Founding stAI tuned & Mentoring Builders

Established a recognized voice in the European AI ecosystem with Gold Impact Writer distinction.

04
Independent Applied AI Stack · 2024 - Present
THE LOCAL-FIRST ODYSSEY

Questioning the Unexamined Cloud Default

Proving that open-weight local execution is a viable architectural tier for modern products.

TRACK RECORD

I also build AI under enterprise constraints.

Local AI experimentation is grounded in experience with production systems, teams and operational trade-offs.

Sky Italia2022 - Present

Data Scientist Manager / Principal AI Engineer & AI Lead Strategy

Directing enterprise AI and data science initiatives across subscriber platforms, content intelligence, and generative AI platforms serving millions of active users.

Independent Local AI Stack2024 - Present

AI Systems Architect & Open-Weight Researcher

Architected and built a comprehensive open-weight Local AI ecosystem spanning reusable inference servers, mobile harnesses, reference applications, and network observability.

stAI tuned2023 - Present

Founder & Managing Director

Founded an independent technical platform delivering deep architectural breakdowns, practical AI roadmaps, and no-hype guides for technical leaders and builders.

WRITING

I build systems and explain what survives contact with reality.

Through stAI tuned, I publish long-form architectural breakdowns designed for practitioners who care about how systems actually work under the hood. No marketing fluff, no buzzwords - just clear mental models, system diagrams, and reproducible implementations.

Explore stAI tuned ↗
01
AI Strategy & Privacy

The Local AI First Principle: Why Defaulting to Cloud is a Risk

Examining why unexamined cloud dependencies introduce privacy, reliability, and vendor risks - and how to establish workload boundaries.

02
Hardware & MLX

Profiling 70B Models on Apple Silicon: Memory, Quantization, and Latency

Comparative benchmarks of 4-bit and 8-bit quantized models on unified memory architectures under sustained production workloads.

03
Agentic Systems

Model Context Protocol (MCP) in Practice: Building Interoperable Agent Workflows

How open protocol standards transform how LLMs interact with local databases, tools, and developer environments without data leakage.

LOCAL AI ADVISORY

Working on the same question?

If you need to decide what should run Local, Hybrid or Cloud, I can apply the same workload-first approach to your architecture.

Evaluate your workload →