Nisharg Nargund
Founder of OpenRAG Innovations and published researcher (IEEE, ACM, Springer, etc.). I build production AI infrastructure - from efficient LLM inference and quantization to RAG systems that work reliably in the real world.
Professional Experience
Researcher — International Institute of Information Technology (IIIT), Hyderabad, India
May 2025 – July 2025
- Implemented BitNet (1-bit LLM) from scratch; evaluated performance under resource-constrained inference on T4 and H100 GPUs.
- Conducted an in-depth architectural analysis of DeepSeek-V3, studying its MoE design, expert routing, and training efficiency innovations.
- Explored edge computing deployment strategies for low-precision LLMs, informing subsequent quantization research (TernaryLM, IJCNN 2026).
AI Strategy Associate — Strategy Boolean(Remote)
Aug 2025 – Sept 2025
- Built scalable GenAI applications including custom RAG chatbots and LLM pipelines for global clients.
- Optimized prompt engineering and model outputs to reduce hallucinations and improve multilingual support.
- Contributed to production-grade solutions across education and enterprise domains.
From Research to Reality: Entrepreneurial Ventures
OpenRAG Innovations Pvt. Ltd.
Founder & Director
OpenRAG is the AI trust layer that verifies outputs before delivery — making LLMs decision-safe for regulated and high-stakes use cases.
Everyone is deploying AI today — but almost no one is measuring how correct those responses are, or how defensible they are when reputational risk is on the line. OpenRAG solves that. Our system lets you query your own data (Excel, DOCX, PDF) through a persona selection layer built for Fintech and Edtech — so you get role-calibrated, context-aware responses instead of the generic outputs every LLM serves by default.
Before any response reaches the user, our system verifies and scores the output pre-delivery — a step almost all companies skip entirely or perform after the fact. DocDynamo isn't a Q&A tool. It's a decision support system that delivers only trusted, persona-backed outputs — giving enterprises a defensible, auditable AI layer they can actually stand behind.
DocDynamo
Multi-agentic backend with a proprietary hallucination removal and verification layer that sits between human and AI. 2,600+ beta users. 2 active B2B contracts.
Conference Partnerships
Official AI partner — ICMEET 2025 (London) and ICDECT 2025 (Bhubaneswar).
CIN: U58201OD2025PTC050816 · Incorporated: 22 Sept 2025 · support@openrag.in
Projects
TernaryLM
Research · Under SubmissionNative ternary-weight language modeling, inspired by the BitNet paradigm
- Designed and implemented a language model architecture using native ternary weights {-1, 0, +1}, inspired by the BitNet paradigm.
- Built custom BitLinear layers, transformer blocks, training infrastructure, and quantization-aware components rather than simply applying post-training quantization.
- Developed experiments around a 132M-parameter model, evaluating language modeling and downstream paraphrase detection performance.
- Investigated the fundamental trade-offs between model memory, computational efficiency, quantization, and quality for resource-constrained inference.
Voice AI Agent
Real-time conversational voice system for production phone calls
- Built a production-style real-time voice agent capable of handling phone conversations through speech recognition, LLM reasoning, and neural text-to-speech.
- Integrated Twilio, WebSockets, STT/TTS services, LiveKit-style real-time infrastructure, and FastAPI to create a low-latency conversational pipeline.
- Implemented streaming audio, call handoffs, interruption handling, session management, and debugging across the entire telephony/audio stack.
- Built versions using multiple voice providers, demonstrating the ability to reason about latency, audio formats, streaming, and vendor abstraction.
GTM Automation Platform
AI-powered sales infrastructure for prospecting and outreach at scale
- Built an end-to-end AI-powered GTM automation platform for identifying, enriching, scoring, and contacting business prospects.
- Engineered pipelines combining company discovery, website scraping, LinkedIn intelligence, Apollo organization search, Hunter domain discovery, and automated enrichment enabling AI-based voice calling outreach to different parts of the world for lead generation.
- Designed the system around FastAPI, PostgreSQL/Supabase, Docker, asynchronous pipelines, state management, and modular enrichment providers.
DocDynamo
2,600+ UsersTrustworthy RAG infrastructure — the core product of OpenRAG Innovations
- An AI infrastructure startup product built around reliable, source-grounded generative AI, evolving from the DocDynamo document intelligence platform.
- Built a full RAG stack spanning document ingestion, retrieval, corrective context generation, web grounding, verification, trust scoring, and reliability analytics.
- The platform evolved from a PDF/YouTube research assistant into a broader de-hallucination and AI trust layer for AI applications.
- Scaled the product to thousands of users while building production APIs, vector-search pipelines, multimodal generation, and an Android/web experience.
AI Research Assistant — ResearchGuru
A precursor to OpenRAG's document intelligence and RAG ecosystem
- Built an AI research assistant designed to help researchers navigate technical literature, documents, and external information.
- Combined document understanding, retrieval, LLM reasoning, summarization, concept extraction, and web information discovery into a unified workflow.
- The system was designed around the practical research workflow rather than a generic chatbot, connecting papers → concepts → explanations → additional sources.
- This work became an important precursor to the broader document intelligence and RAG product ecosystem at OpenRAG.
Research Papers
TernaryLM: Memory-Efficient Language Modeling via Native 1.5-Bit Quantization with Adaptive Layer-wise Scaling
Status: Accepted · Camera Ready
International Joint Conference on Neural Networks, Maastricht, Netherlands, 2026
Demonstrates native 1.5-bit quantization with per-layer precision scaling — enabling competitive LLM performance at extreme memory compression for edge deployment.
Optimization of Latent-Space Compression using Game-Theoretic Techniques for Transformer-Based Vector Search
Status: Accepted
A Transformer-Based Model for Enhanced Tokenization and Generation in Hindi Natural Language Processing
Status: Presented (Pending Publication)
Neural Orchestration for Multi-Agent Systems: A Deep Learning Framework for Optimal Agent Selection
Status: Presented (Pending Publication)
Model Context Protocol (MCP): A Lightweight, Modular Framework for Tool-Augmented LLM Agents
Status: Accepted (Camera Ready)
From Chalkboards to Chatbots: Exploring Teacher and Student Readiness for Agentic AI in Indian Schools
Status: Presented (Pending Publication)
Conversational Text Extraction with Large Language Models Using Retrieval-Augmented Systems
Status: Published
Innovative Fusion of LSTM and Bi-GRU Networks for Enhanced Hate Speech Detection in Social Media
Status: Published
Deep Learning in Industry 4.0: Transforming Manufacturing through Data-Driven Innovation
Status: Published
Advancements in Computer Vision and Machine Learning for Food Quality Evaluation: A Comprehensive Review for the Food Industry
Status: Published
Articles
2024
Beyond Reality: Virtual Worlds Shaping the Future of Food Science
Food Infotech Magazine — January 2024
Read Article →Advancements in Food Safety and Quality Evaluation Using Computer Vision and Machine Learning
Food, Marketing and Technology Magazine — March 2024
Read Article →Implementing the Technological Intelligence in Agricultural Produce
Food, Marketing and Technology Magazine — June 2024
Read Article →Edible and Biodegradable Packaging Innovations
Food, Marketing and Technology Magazine — December 2024
Read Article →2023
Two Layer Classifiers of MVS for Effective Grading, Safety & Quality Evaluation in Food Industry
Food Infotech Magazine — April 2023
Read Article →Crunching the number of safer foods: how big data is transforming food safety
Food Infotech Magazine — May 2023
Read Article →Cracking the Code! Unveiling the Hidden World of Food Safety with MicroWaves & ML
Food Infotech Magazine — June 2023
Read Article →Revolutionizing Agri-Food Supply Chain: Harnessing the Power of IoT
Food Infotech Magazine — July 2023
Read Article →Revolutionizing Food Safety: How BlockChain and Lighting Network are changing the game
Food Infotech Magazine — September 2023
Read Article →Revolutionizing Brain Tumor Detection: The CNN-based Medical Imaging Breakthrough
Informs London Magazine — December 2023
Read Article →Blog — Writing on Medium
My side hustle and running passion for AI — deep dives, model breakdowns, and industry analysis, published regularly on Medium.
August 15, 2026
Understanding Latency in Multi Agentic AI Systems | DeepDive #1
Model latency is only 30–40% of the total latency shown by the system.
Read on Medium →
July 21, 2026
Kimi K3: 2.8T Most Capable Model by Moonshot.ai
Moonshot's flagship 2.8T-parameter model — novel attention and MoE design for efficiency.
Read on Medium →
June 9, 2026
The Complete Guide to Attention Variants in Transformers: From Scaled Dot-Product to Flash
A comprehensive overview of transformer attention mechanisms and their evolution.
Read on Towards AI →
May 5, 2026
The Hidden Engine Behind Fast LLM Inference | Why It's Necessary for All Wrapper Applications
Infrastructure optimizations powering rapid LLM responses beyond model speed alone.
Read on Medium →
April 19, 2026
The Quiet Revolution Happening Inside Your Software — and Why Most Developers Are Already Behind
How agentic AI systems are quietly reshaping software development.
Read on Medium →
April 11, 2026
How Perplexity AI Makes Money on Top of GPT, Claude, Gemini: Inside the Answer Engine Business
Analyzing a $21B startup's business model built on third-party LLMs.
Read on Medium →
April 2, 2026
Meta's TRIBE v2: The AI That Predicts How Your Brain Reacts to the World
A neuroscience foundation model trained on 500+ hours of fMRI data.
Read on Medium →
February 8, 2026
Kimi K2.5: The Next Milestone in Multimodal and Agentic AI
Moonshot's multimodal model advancing visual understanding and agentic capability.
Read on Medium →
January 25, 2026
Why LLMs Forget Long Conversations | AI Replacing Humans Narrative
Context window limits keep LLMs from developing genuine memory, unlike human cognition.
Read on Medium →
January 5, 2026
Compression Is Easy. Preserving Meaning Is Not | Compressed Language Models — Beginner 101
Shrinking the model is the easy part. Keeping it intelligent? That's hard.
Read on Medium →Authored Books
Fundamentals of Convolutional Neural Networks with TensorFlow
A deep dive into Computer Vision, Object Detection, AI, CNNs, and TensorFlow.
Available on: Amazon, Kindle, Flipkart
Get The BookTrust in the Age of Agentic AI Economy
A C-suite guide to building trustworthy, reliable agentic AI systems at scale. Drawn from the infrastructure OpenRAG Innovations has built in production — covering hallucination mitigation, retrieval reliability, multi-agent coordination, and the emerging trust frameworks enterprises need as AI systems become autonomous decision-makers.
Co-authored under OpenRAG Innovations Pvt. Ltd.