LOCAL-FIRST AI & PRIVACY FOUNDATIONS

Building the infrastructure
for privacy-first AI products.

I build the foundations that let AI products work with sensitive data while keeping privacy, control, and critical processing close to the user.

Current work spans local LLM inference, private speech, and on-device AI—validated through products that protect sensitive conversations, documents, and financial data.

Applications such as ClosedRoom, RedactGuard, and Aura Finance are built on local LLM, speech recognition, and Android infrastructure running on private user-owned devices, with no cloud required.
  1. Applications: ClosedRoom, RedactGuard, and Aura Finance.
  2. Infrastructure: Local LLM Server, Local ASR Server, and Android LLM Harness.
  3. Hardware: laptop, desktop, and Android devices, with no cloud required.

VISION

AI products should be able to understand sensitive information without forcing people to surrender control of their conversations, documents, or personal data.

MISSION

Build reusable local AI and data-protection foundations that make privacy-first products possible across desktop and mobile.

  • Local by default
  • Privacy by architecture
  • Open & auditable

FROM LOCAL MODELS TO REAL PRODUCTS

Running a model locally is only the beginning.

Shipping product-grade, privacy-first AI requires far more than weights on a device. It demands solving four key technical challenges:

01

Runtime lifecycle

Model management, hardware compatibility, dynamic switching & zero-downtime updates.

02

Observability & telemetry

Performance tracking, latency metrics, accuracy signals & memory profiling without privacy leakage.

03

Hardware adaptation

RAM/VRAM constraints, resident memory optimization & thermal management on desktop and mobile.

04

Developer ergonomics

Unified OpenAI-compatible APIs, reliable local IPC & intuitive user experiences.

THE INFRASTRUCTURE STACK (THREE PILLARS)

Three infrastructure pillars

Each project solves a different part of the local-AI product stack. Reusable infrastructure, not isolated experiments.

ACTIVE

Local LLM Server

High-performance local inference server with model lifecycle, dynamic routing and telemetry.

RustgRPCQuantizationTooling
Explore the pillar

Provides an OpenAI-compatible API across local backends such as GGUF and Apple Silicon MLX, managing resident runtimes, configuration and telemetry.

  • OpenAI-compatible local API
  • GGUF and Apple Silicon (MLX) backends
  • Runtime model switching & multi-residency
ACTIVE

Local ASR Server

Privacy-first speech recognition server optimized for sub-second latency and accuracy.

WhisperVADStreamingDiarization
Explore the pillar

Handles microphone & system audio capture, local session storage, and Whisper transcription on Apple Silicon with zero cloud calls.

  • Microphone & system-audio capture
  • Local Whisper transcription with VAD
  • OpenAI-compatible audio endpoint
IN PROGRESS

Android Local LLM Harness

On-device LLM runtime harness for Android apps with strict resource control.

KotlinML RuntimeMemoryOffline
Explore the pillar

Reusable runtime for embedding GGUF models in native Kotlin and Capacitor Android applications with lifecycle control and diagnostics.

  • GGUF import & cryptographic integrity check
  • Managed model and context lifecycle
  • Real-time streaming & cancellation

ARCHITECTURE RELATIONSHIP

Local LLM ServerLocal ASR ServerAndroid LLM Harness
APIs / Protocols (gRPC, WebSocket, Local IPC)Reference Applications (ClosedRoom, RedactGuard, Aura Finance)

PROVING GROUNDS

Reference applications that validate the stack

ClosedRoom, RedactGuard, and Aura Finance demonstrate how privacy-first AI protects conversations, sensitive documents, and personal financial data in practice.

ClosedRoom logo
ACTIVE

ClosedRoom

Privacy-first meeting intelligence.

Local-first meeting capture, Whisper transcription and structured insights. Your conversations stay 100% on-device.

TranscriptionSummariesAction ItemsSearch
What it validates
  • Local ASR & private speech recognition
  • Local LLM analysis & structured insights
  • Installable desktop app packaging & offline workflows
RedactGuard logo
ACTIVE

RedactGuard

Local document anonymization & PII protection.

A local document-anonymisation workflow that detects PII via local LLMs, provides human-in-the-loop review, and selectively redacts sensitive data. Data minimisation built into the architecture.

Doc IntelligencePII DetectionHuman ReviewSelective RedactionSafe Export
What it validates
  • Local document intelligence & PII detection
  • Human review & selective AI redaction
  • Data minimisation & safe document portability
Aura Finance logo
IN PROGRESS

Aura Finance

Privacy-first personal finance.

Local AI for transaction categorization, forecasting and budget planning. Your financial data never leaves your device.

Spending InsightsForecastingBudgetsPrivacy First
What it validates
  • Privacy-first product design & local data model
  • Mobile UX & local storage sovereignty
  • Proving ground for Android on-device AI integration
Reference applications validate the infrastructure and show what it can enable. They are the beginning, not the limit, of the ecosystem.

WHAT THIS ENABLES

What the stack makes easier to build next

Exploration areas unlocked by core local AI infrastructure.

01

Meeting intelligence

Private, local summaries & action items from audio.

02

Personal finance

Local transaction analysis, budget forecasting & planning.

03

Document analysis

Query confidential documents without cloud exposure.

04

Professional assistants

Domain-aware copilots that keep workflows private.

05

Personal knowledge

Your notes, custom models, and data sovereignty.

06

Offline productivity

Work anywhere without sacrificing privacy.

07

Sensitive-data workflows

Process confidential data safely on local hardware.

08

Mobile AI applications

Powerful on-device experiences, completely offline-first.

GUIDING ARCHITECTURAL PRINCIPLES

Privacy-first is an architectural discipline

Core tenets guiding engineering decisions across infrastructure and application layers.

01

Local by default

Compute happens on your device, by default.

02

Privacy by architecture

Privacy is designed in, not bolted on.

03

Observable & measurable

We instrument, benchmark and share results.

04

Reusable across apps

Build once. Use everywhere. Composability first.

05

Evidence over claims

We ship, test and publish evidence.

REAL-WORLD EVIDENCE & TRACK RECORD

Proven capability to deliver at scale

"My local-first work is independent, but it builds on years spent designing, evaluating, and scaling AI systems in complex organisations."

— Daniele Moltisanti
Enterprise
Scale & Governance

AI Technical Leadership

Experience leading AI initiatives in complex organisations, turning experiments into adopted, governed, and scaled systems.

80+
Research & Insights

80+ Articles Published

Long-form research, model evaluations, and practical LLM insights shared with the technical community.

150+
Engineering Mentorship

150+ AI Projects Reviewed

Hands-on feedback, architecture guidance, and mentoring for emerging AI engineers and technical teams.

Top 111
Award & Recognition

Nova 111 (2022)

Selected by Nova Talent among emerging Italian talent driving innovation and technology leadership.

ABOUT DANIELE

I work where AI infrastructure, product strategy, and clear engineering intersect.

I am Daniele Moltisanti, an AI technical leader and builder based in Milan. My work combines technical architecture, product thinking, and the ability to transform complex AI research into understandable decisions and reliable, intuitive software.

Alongside my independent work, I lead and design AI initiatives in complex organisations, translating technical possibilities into systems that can be adopted, governed, and scaled.

I also write, teach, and speak about AI architecture, strategy, and the path from experimentation to reliable products.

Daniele Moltisanti

COLLABORATE

Let's build the future of local AI infrastructure.

Open to collaboration, architecture feedback, and engineering challenges across local LLM performance, Android on-device inference, and privacy-first product design.

danielemoltisanti@gmail.com Milan, Italy