AI Engineer & Builder

Local-first LLM systems, shipped end-to-end

I design, build, and ship AI that runs on the hardware people own — offline, private, and in the languages the industry ignores.

About Me

I'm an AI engineer and consultant working across the full AI lifecycle — from evaluating model architectures and benchmarking performance, to building production-ready knowledge systems and deploying private AI infrastructure.

My work turns emerging AI research into practical business solutions: intelligent retrieval systems that answer real questions, agentic workflows that automate real operations, and local-first inference that keeps organizational data exactly where it belongs — under the organization's control. I've optimized inference across llama.cpp, GGUF, Ollama, and vLLM, and built platforms that make AI deployment simple enough for non-technical teams to adopt.

Before AI, I founded and led a peer-to-peer trading operation — solving real problems under real constraints. I bring that standard to every system I build: evaluate the trade-offs honestly, choose the architecture deliberately, and ship something that survives contact with production.

Full-Lifecycle AI Delivery

Architecture evaluation, RAG systems, agentic workflows, and knowledge management — designed, built, and deployed end-to-end.

Intelligent Retrieval & Knowledge Systems

Business-grade RAG built on multiple chunking, retrieval, reranking, and knowledge-grounding strategies — ChromaDB, FAISS, and vector search at scale.

Private & Local AI Infrastructure

On-premise, sovereign AI deployment with inference optimized across llama.cpp, GGUF, Ollama, and vLLM — data never leaves your environment.

Evaluation-Driven Engineering

Every system benchmarked on accuracy, latency, scalability, resource utilization, and operational fit. "Works in a demo" is not a standard I ship to.

2024 – Present Independent AI Consultant
2018 – 2021 Technical Lead, P2P Crypto — Arta Services
Languages Persian (native) · English · Pashto · Hindi

Core Competencies

Hands-on AI engineering across the complete lifecycle — model research and benchmarking, retrieval and agentic systems, private infrastructure, and the products that run on it.

LLM & Applied AI Engineering

Shipped production LLM applications end-to-end — model fine-tuning, quantization, and CPU-efficient inference on llama.cpp, Ollama, and raw PyTorch, built for the hardware most users actually own.

RAG, Agents & Knowledge Systems

Architected retrieval and agentic systems on LangChain and LangGraph — hybrid BM25 + semantic search, reranking, and citation-grounded answers that can be verified, not just trusted.

Desktop & Mobile Product Development

Delivered cross-platform AI products for non-technical users — PySide6 desktop applications and Kotlin/Android assistants with zero-CLI, on-device experiences.

Multilingual & Low-Resource NLP

Built AI for underrepresented languages — Pashto, Dari, Persian, Urdu — with custom glossaries, context-aware translation, and evaluation that goes beyond English benchmarks.

Privacy-First & Sovereign AI Architecture

Designed offline-first systems where data never leaves the device — zero telemetry, on-prem deployment, and compliance alignment for legal and regulated industries.

SEO · GEO · AEO & Digital Growth

Own search and AI discoverability end-to-end for every product I ship — technical SEO, structured data, and content engineered to rank, and to be cited, by search, generative, and answer engines.

Engineering Approach

Complex Systems, Simple Interfaces

Hard problems deserve boring, reliable solutions. I pair deep technical work with interfaces that non-technical users can actually operate — no CLI required.

Tech: PySide6, modular plugin architecture

Underrepresented Languages, First-Class

Most AI ignores most of the world's languages. I build context-aware pipelines for Pashto, Dari, Persian, and Urdu — with glossaries, culture, and evaluation that go beyond English.

Tech: LangChain, Mistral, GGUF

Performance on Real Hardware

I optimize for the machines people actually own — quantization, memory efficiency, and inference tuning that turn ordinary laptops into capable AI workstations.

Tech: llama.cpp, quantization, GGUF

Privacy as Architecture, Not Policy

Data protection is designed in from the start — offline-first processing, zero telemetry, and compliance alignment for legal and regulated industries.

Tech: Local inference, GDPR alignment

Flagship Project: Lawyer Assistant

Lawyer Assistant — private AI legal research app screenshot

Private AI legal research for your own documents — ask in plain English, get cited answers, 100% on your machine.

Challenges Addressed

  • Confidentiality: Cloud AI tools leak sensitive legal data and break attorney–client privilege
  • Verification Gap: Generic chatbots answer from memory without showing their sources
  • Document Overload: Reviewing hundreds of pages of contracts and case files is slow and error-prone

Key Differentiators

Answers with Citations

Every claim grounded in retrieved passages with inline source links

Compliance Playbook Scan

Custom rules flag, rate, and explain violations in any document

Hybrid Search Engine

Semantic + keyword (BM25) retrieval fused and re-ranked

100% Private & Offline

Windows, macOS & Linux — no account, no cloud, no telemetry

Lawyer Assistant Performance

Citation Accuracy
96%
Data Privacy
100%
Search Recall
94%
Setup Time Reduction
90%

Visit Website

Flagship Project: GPT Calendar

GPT Calendar — calendar screen
Calendar
GPT Calendar — finance tracking screen
Finance
GPT Calendar — location-based reminder screen
Location Alerts
GPT Calendar — planner screen
Planner
GPT Calendar — AI assistant screen
AI Assistant

The voice-first AI personal assistant that replaces 5 apps — calendar, reminders, finance tracking, location alerts, and tasks in one Android app.

Challenges Addressed

  • App Fragmentation: Users juggle 5+ separate apps for calendar, reminders, expenses, and location alerts
  • Slow Input: Typing every task and reminder is tedious — voice is far faster
  • Manual Finance Tracking: Nobody wants to log every expense by hand
  • Privacy: Most assistants phone home; this one offers fully on-device AI

Key Differentiators

Voice-First Interface

"Kiro" wake word + voice commands for hands-free use

Smart Finance Tracking

Auto-parses SMS transactions into your expense log

Location-Based Reminders

Geofencing alerts that fire where you actually are

On-Device AI

Works offline with privacy-first local intelligence

Smart Calendar Performance

Voice Command Accuracy
93%
SMS Parsing Accuracy
96%
On-Device Privacy
100%
Apps Replaced
5 → 1

Visit Website

Featured Platform: GGUFLoader

Transform any laptop into a secure, customizable, multilingual AI workstation.

Challenges Addressed

  • Steep Learning Curves: Local LLM deployment typically demands CLI + manual config
  • Rigid Tooling: Hard to adapt tools to specialized education, research, or enterprise use
  • Hardware Uncertainty: Users worry about GPU/CPU compatibility

Key Differentiators

Zero-CLI UI

Fully graphical interface with drag & drop model loading

Real-Time Dashboard

Monitor RAM, VRAM, and active threads

Extensible Plugins

Build custom chat UIs, translators, document processors

Multilingual Modules

Ready for Pashto, Dari, Persian educational pipelines

GGUF Loader Performance

Setup Time Reduction
85%
Memory Efficiency
78%
Multilingual Accuracy
92%
Plugin Compatibility
95%

Programs & Applications I Built

Mobile AI Assistant

Problem: AI assistants are too heavy for everyday phones

Solution: Lightweight, mobile-optimized AI assistant built in Kotlin for on-device use

raw-pytorch-minigpt

Problem: Frameworks hide how transformers actually work

Solution: GPT-style LLM built from scratch in raw PyTorch — custom BPE tokenizer, decoder-only transformer, SwiGLU ablation, two-stage fine-tuning

LLM-Toolkit

Focus: Open-source Python toolkit (MIT) for local LLM workflows and model tooling

Offline AI Assistant

Private

Problem: Internet dependence, limited support for underrepresented languages

Solution: Hybrid Mistral + llama.cpp architecture for offline Persian/Urdu dialogue

SEO · GEO · AEO — Websites I Handle & Optimize

Part 2 — Beyond building software, I own the online presence of every product I ship. I handle end-to-end SEO (search engine optimization), GEO (generative engine optimization), and AEO (answer engine optimization): technical SEO, structured data & JSON-LD, FAQ schemas, sitemaps, canonical URLs, Open Graph/Twitter cards, content strategy, and making each site get discovered — and cited — by search engines and AI assistants.

GGUF Loader

SEO · GEO · AEO: Product site with model downloads, docs, and local AI content targeting high-intent search queries.

Local AI Zone

SEO · GEO · AEO: Daily-updated model discovery site with expert guides, FAQ content, and structured data built to rank for model queries.

Haal Lab

SEO · GEO · AEO: Enterprise AI infrastructure site — private RAG, agents, and sovereign AI positioning for regulated industries.

Lawyer Assistant

SEO · GEO · AEO: FAQ-driven content, download pages, and citation-friendly product copy that AI assistants can quote.

Hussain Nazary Portfolio

SEO · GEO · AEO: Person & FAQPage structured data, canonical URLs, sitemap, and blog posts optimized to rank and get cited by AI engines.

Unified Architecture

GGUFLoader Core Engine
Addon Plugins
Lawyer Assistant
Multilingual Assistant
PySide6 Interface
Chat UI
Document Processor
Translator Module
Model Integration
llama.cpp
GGUF Models
LangChain Pipelines
All components run entirely offline with no external dependencies

Why Local-First AI

Privacy as Architecture

Your data never leaves your device — no telemetry, no third-party processing, no ambiguity about who sees client material. Privacy is designed in, not added as a policy.

Optimized for Real Hardware

Built for the machines you have, not the ones a vendor wishes you'd buy — quantized models, efficient CPU inference, and memory tuning that keep costs low and answers fast.

Multilingual by Default

Most of the world's languages are underserved by mainstream AI. I build for Pashto, Dari, Persian, and Urdu with the same rigor as English — glossaries, context, and evaluation included.

Latest Blog Posts

LFM Models: Liquid Foundation Models Explained (2026 Guide)

Liquid AI's non-transformer family — LFM-40B down to the tiny LFM2.5 on-device line, with benchmarks and how to run them.

Read more

Ant Lab AI Models: The Ling Family Behind Ant Group (2026 Guide)

InclusionAI's Ling, Ring, and Ming families — from the 1T flagship to the agent-focused Ling 3.0 Flash, with benchmarks.

Read more

Top 10 AI Models Under 12B for Coding in 2026 (Ranked by Benchmarks)

Qwen3.5-9B, Gemma 4 12B, Phi-4-mini, Yi-Coder 9B — HumanEval scores and a head-to-head table for local coding.

Read more

Top 10 AI Models Under 12B Parameters in 2026 (Ranked by Power)

Gemma 4 12B, Qwen3.5-9B, Phi-4-mini, and more — the newest powerful models you can actually run, with sizes and hardware.

Read more

Top 10 RAG Tools in 2026 (Ranked by Use Case)

LlamaIndex, LangChain, Haystack, RAGFlow, Dify, txtai, RAGAS, and more — ranked by use case with a comparison table.

Read more

Top 10 Vector Databases in 2026 (Ranked by Use Case)

ChromaDB, FAISS, Qdrant, pgvector, Weaviate, Milvus, Pinecone, and more — compared by scale, hosting, and use case.

Read more

Top 10 Embedding Models for RAG in 2026 (Ranked by MTEB Score)

Qwen3-Embedding, BGE-M3, Gemini, Voyage, Cohere, OpenAI, mxbai, and Nomic — with MTEB scores and a decision guide.

Read more

Top 10 Open Source AI Models in 2026 (Ranked by Capability & Performance)

Kimi K3, GLM-5.2, DeepSeek-V4, Qwen 3.5, GPT-OSS, and Gemma 4 — benchmark scores, licenses, and what you can actually run.

Read more

Top 10 GGUF Models Ranked by RAM Size (Download Guide)

What to download for 8GB, 16GB, 24GB, and 64GB machines — with exact Q4_K_M file sizes.

Read more

Top 10 AI Models You Can Run Locally in 2026 (Ranked by Hardware)

Qwen3, GPT-OSS, GLM, Gemma, DeepSeek, and more — ranked by use case and the hardware you actually own.

Read more

Offline AI for Regulated Industries: Legal, Health, Finance

Privilege, HIPAA, and SOX — what offline AI genuinely solves for regulated teams, and what it doesn't.

Read more

Can You Run a Private LLM for Your Business? The Honest Guide

Real costs, when on-premise beats cloud APIs, what you give up, and how to start without overcommitting.

Read more

Embeddings Explained: How Semantic Search Actually Works

What embeddings are, how transformer models turn text into vectors, and how cosine similarity powers semantic search.

Read more

Reranking in RAG: Why Retrieval Order Matters

How cross-encoders reorder retrieval results for better answers — with code to add reranking to your RAG pipeline.

Read more

ChromaDB vs FAISS: Choosing a Vector Database for RAG

Database vs search library — persistence, metadata filtering, indexing, and scale, with migration guidance.

Read more

What Is Hybrid Search? BM25 + Semantic Retrieval Explained

How keyword and vector retrieval complement each other, how Reciprocal Rank Fusion combines them, and why it improves RAG answers.

Read more

How to Build a RAG System in 30 Minutes (Local, Free)

Chunk, embed, retrieve, and answer with ChromaDB and Ollama — a fully local RAG pipeline, no cloud or API keys.

Read more

How to Convert a Hugging Face Model to GGUF (Step by Step)

Convert any supported Hugging Face model to GGUF with llama.cpp's convert_hf_to_gguf.py and llama-quantize — no GPU needed.

Read more

8 Best Local LLM Tools in 2026 (Ranked by Use Case)

Ollama, LM Studio, llama.cpp, GGUF Loader, Jan, GPT4All, text-generation-webui, and vLLM — ranked by use case with a comparison table.

Read more

Q4_K_M vs Q8_0: Which Quantization Should You Download?

Real numbers on file size, memory, and quality loss — plus a decision guide for picking the right GGUF quant.

Read more

Ollama vs llama.cpp: Which Should You Use for Local AI in 2026?

Managed runtime vs raw engine — interface, API, quantization control, and the honest performance numbers, with clear recommendations.

Read more

What is GGUF? The File Format Behind Local AI Explained

What GGUF actually is, why it replaced GGML and became the standard for local LLMs, and how quantization really works.

Read more

How to Run GGUF Models Locally in 2026: The Complete Beginner's Guide

Set up Ollama, llama.cpp, LM Studio, or GGUF Loader on your own machine — with model picks for 8–16GB RAM and quantization explained.

Read more

Lawyer Assistant: A Privacy-First Legal AI Built on a Local RAG Pipeline

How a fully local RAG pipeline grounds legal answers in your own documents — with citations, compliance scans, and zero cloud.

Read more

AI IDE Comparison (Feb 2026): Copilot vs Cursor vs JetBrains AI vs Kiro + Antigravity + Codex

Updated comparison of pricing, limits, and model access across top AI coding tools.

Read more

Latest AI Model Updates (Feb 2026): GPT-5.3-Codex, Claude Opus 4.6, Gemini 3 Deep Think

Updated Feb 11, 2026 with verified releases, rollouts, and official source links.

Read more

Getting Started with GGUF Models

A comprehensive guide to converting and running quantized models with llama.cpp for optimal performance.

Read more

Get In Touch

Email

hussainnazary475@gmail.com

Open Source

Explore my projects on GitHub