Runtime lifecycle
Model management, hardware compatibility, dynamic switching & zero-downtime updates.
I build the foundations that let AI products work with sensitive data while keeping privacy, control, and critical processing close to the user.
Current work spans local LLM inference, private speech, and on-device AI—validated through products that protect sensitive conversations, documents, and financial data.
VISION
MISSION
Build reusable local AI and data-protection foundations that make privacy-first products possible across desktop and mobile.
FROM LOCAL MODELS TO REAL PRODUCTS
Shipping product-grade, privacy-first AI requires far more than weights on a device. It demands solving four key technical challenges:
Model management, hardware compatibility, dynamic switching & zero-downtime updates.
Performance tracking, latency metrics, accuracy signals & memory profiling without privacy leakage.
RAM/VRAM constraints, resident memory optimization & thermal management on desktop and mobile.
Unified OpenAI-compatible APIs, reliable local IPC & intuitive user experiences.
THE INFRASTRUCTURE STACK (THREE PILLARS)
Each project solves a different part of the local-AI product stack. Reusable infrastructure, not isolated experiments.
High-performance local inference server with model lifecycle, dynamic routing and telemetry.
Provides an OpenAI-compatible API across local backends such as GGUF and Apple Silicon MLX, managing resident runtimes, configuration and telemetry.
Privacy-first speech recognition server optimized for sub-second latency and accuracy.
Handles microphone & system audio capture, local session storage, and Whisper transcription on Apple Silicon with zero cloud calls.
On-device LLM runtime harness for Android apps with strict resource control.
Reusable runtime for embedding GGUF models in native Kotlin and Capacitor Android applications with lifecycle control and diagnostics.
ARCHITECTURE RELATIONSHIP
PROVING GROUNDS
ClosedRoom, RedactGuard, and Aura Finance demonstrate how privacy-first AI protects conversations, sensitive documents, and personal financial data in practice.

Privacy-first meeting intelligence.
Local-first meeting capture, Whisper transcription and structured insights. Your conversations stay 100% on-device.

Local document anonymization & PII protection.
A local document-anonymisation workflow that detects PII via local LLMs, provides human-in-the-loop review, and selectively redacts sensitive data. Data minimisation built into the architecture.

Privacy-first personal finance.
Local AI for transaction categorization, forecasting and budget planning. Your financial data never leaves your device.
WHAT THIS ENABLES
Exploration areas unlocked by core local AI infrastructure.
Private, local summaries & action items from audio.
Local transaction analysis, budget forecasting & planning.
Query confidential documents without cloud exposure.
Domain-aware copilots that keep workflows private.
Your notes, custom models, and data sovereignty.
Work anywhere without sacrificing privacy.
Process confidential data safely on local hardware.
Powerful on-device experiences, completely offline-first.
GUIDING ARCHITECTURAL PRINCIPLES
Core tenets guiding engineering decisions across infrastructure and application layers.
Compute happens on your device, by default.
Privacy is designed in, not bolted on.
We instrument, benchmark and share results.
Build once. Use everywhere. Composability first.
We ship, test and publish evidence.
REAL-WORLD EVIDENCE & TRACK RECORD
"My local-first work is independent, but it builds on years spent designing, evaluating, and scaling AI systems in complex organisations."
Experience leading AI initiatives in complex organisations, turning experiments into adopted, governed, and scaled systems.
Long-form research, model evaluations, and practical LLM insights shared with the technical community.
Hands-on feedback, architecture guidance, and mentoring for emerging AI engineers and technical teams.
Selected by Nova Talent among emerging Italian talent driving innovation and technology leadership.
KNOWLEDGE & PUBLIC RESEARCH
Continuous synthesis across three public channels: Build → Document → Distribute.
ABOUT DANIELE
I am Daniele Moltisanti, an AI technical leader and builder based in Milan. My work combines technical architecture, product thinking, and the ability to transform complex AI research into understandable decisions and reliable, intuitive software.
Alongside my independent work, I lead and design AI initiatives in complex organisations, translating technical possibilities into systems that can be adopted, governed, and scaled.
I also write, teach, and speak about AI architecture, strategy, and the path from experimentation to reliable products.

COLLABORATE
Open to collaboration, architecture feedback, and engineering challenges across local LLM performance, Android on-device inference, and privacy-first product design.