Nisharg Nargund

Nisharg Nargund

Founder of OpenRAG Innovations and published researcher (IEEE, ACM, Springer, etc.). I build production AI infrastructure - from efficient LLM inference and quantization to RAG systems that work reliably in the real world.

Currently seeking MS Research / PhD positions in efficient AI & LLM systems

Professional Experience

Researcher — International Institute of Information Technology (IIIT), Hyderabad, India

May 2025 – July 2025

  • Implemented BitNet (1-bit LLM) from scratch; evaluated performance under resource-constrained inference on T4 and H100 GPUs.
  • Conducted an in-depth architectural analysis of DeepSeek-V3, studying its MoE design, expert routing, and training efficiency innovations.
  • Explored edge computing deployment strategies for low-precision LLMs, informing subsequent quantization research (TernaryLM, IJCNN 2026).

AI Strategy Associate — Strategy Boolean(Remote)

Aug 2025 – Sept 2025

  • Built scalable GenAI applications including custom RAG chatbots and LLM pipelines for global clients.
  • Optimized prompt engineering and model outputs to reduce hallucinations and improve multilingual support.
  • Contributed to production-grade solutions across education and enterprise domains.

From Research to Reality: Entrepreneurial Ventures

OpenRAG Innovations Pvt. Ltd.

Founder & Director

2,600+ Users 2 B2B Clients Bhubaneswar · India

OpenRAG is the AI trust layer that verifies outputs before delivery — making LLMs decision-safe for regulated and high-stakes use cases.

Everyone is deploying AI today — but almost no one is measuring how correct those responses are, or how defensible they are when reputational risk is on the line. OpenRAG solves that. Our system lets you query your own data (Excel, DOCX, PDF) through a persona selection layer built for Fintech and Edtech — so you get role-calibrated, context-aware responses instead of the generic outputs every LLM serves by default.

Before any response reaches the user, our system verifies and scores the output pre-delivery — a step almost all companies skip entirely or perform after the fact. DocDynamo isn't a Q&A tool. It's a decision support system that delivers only trusted, persona-backed outputs — giving enterprises a defensible, auditable AI layer they can actually stand behind.

DocDynamo

Multi-agentic backend with a proprietary hallucination removal and verification layer that sits between human and AI. 2,600+ beta users. 2 active B2B contracts.

Conference Partnerships

Official AI partner — ICMEET 2025 (London) and ICDECT 2025 (Bhubaneswar).

Explore OpenRAG →

CIN: U58201OD2025PTC050816  ·  Incorporated: 22 Sept 2025  ·  support@openrag.in

Projects

1

TernaryLM

Research · Under Submission

Native ternary-weight language modeling, inspired by the BitNet paradigm

  • Designed and implemented a language model architecture using native ternary weights {-1, 0, +1}, inspired by the BitNet paradigm.
  • Built custom BitLinear layers, transformer blocks, training infrastructure, and quantization-aware components rather than simply applying post-training quantization.
  • Developed experiments around a 132M-parameter model, evaluating language modeling and downstream paraphrase detection performance.
  • Investigated the fundamental trade-offs between model memory, computational efficiency, quantization, and quality for resource-constrained inference.
BitLinear Ternary Quantization Transformers PyTorch Edge Inference
View on Hugging Face →
2

Voice AI Agent

Real-time conversational voice system for production phone calls

  • Built a production-style real-time voice agent capable of handling phone conversations through speech recognition, LLM reasoning, and neural text-to-speech.
  • Integrated Twilio, WebSockets, STT/TTS services, LiveKit-style real-time infrastructure, and FastAPI to create a low-latency conversational pipeline.
  • Implemented streaming audio, call handoffs, interruption handling, session management, and debugging across the entire telephony/audio stack.
  • Built versions using multiple voice providers, demonstrating the ability to reason about latency, audio formats, streaming, and vendor abstraction.
Twilio WebSockets STT / TTS LiveKit FastAPI Streaming Audio
3

GTM Automation Platform

AI-powered sales infrastructure for prospecting and outreach at scale

  • Built an end-to-end AI-powered GTM automation platform for identifying, enriching, scoring, and contacting business prospects.
  • Engineered pipelines combining company discovery, website scraping, LinkedIn intelligence, Apollo organization search, Hunter domain discovery, and automated enrichment enabling AI-based voice calling outreach to different parts of the world for lead generation.
  • Designed the system around FastAPI, PostgreSQL/Supabase, Docker, asynchronous pipelines, state management, and modular enrichment providers.
FastAPI PostgreSQL / Supabase Docker Apollo / Hunter APIs Async Pipelines AI Voice Outreach
4

DocDynamo

2,600+ Users

Trustworthy RAG infrastructure — the core product of OpenRAG Innovations

  • An AI infrastructure startup product built around reliable, source-grounded generative AI, evolving from the DocDynamo document intelligence platform.
  • Built a full RAG stack spanning document ingestion, retrieval, corrective context generation, web grounding, verification, trust scoring, and reliability analytics.
  • The platform evolved from a PDF/YouTube research assistant into a broader de-hallucination and AI trust layer for AI applications.
  • Scaled the product to thousands of users while building production APIs, vector-search pipelines, multimodal generation, and an Android/web experience.
RAG Vector Search Trust Scoring Multimodal Generation Android / Web
Explore OpenRAG →
5

AI Research Assistant — ResearchGuru

A precursor to OpenRAG's document intelligence and RAG ecosystem

  • Built an AI research assistant designed to help researchers navigate technical literature, documents, and external information.
  • Combined document understanding, retrieval, LLM reasoning, summarization, concept extraction, and web information discovery into a unified workflow.
  • The system was designed around the practical research workflow rather than a generic chatbot, connecting papers → concepts → explanations → additional sources.
  • This work became an important precursor to the broader document intelligence and RAG product ecosystem at OpenRAG.
Document Understanding Retrieval Summarization Concept Extraction

Research Papers

Springer · ICDCIT 2026 📍 KIIT University

Optimization of Latent-Space Compression using Game-Theoretic Techniques for Transformer-Based Vector Search

Status: Accepted

Read on arXiv →
Springer · FICTA 2025 📍 London, UK

A Transformer-Based Model for Enhanced Tokenization and Generation in Hindi Natural Language Processing

Status: Presented (Pending Publication)

Link TBA Best Paper Award
Springer · PReMI 2025 📍 IIT Delhi

Neural Orchestration for Multi-Agent Systems: A Deep Learning Framework for Optimal Agent Selection

Status: Presented (Pending Publication)

Read on arXiv →
IEEE · ISED 2025 📍 NIT Raipur

Model Context Protocol (MCP): A Lightweight, Modular Framework for Tool-Augmented LLM Agents

Status: Accepted (Camera Ready)

Link TBA
ACM · COMPUTE 2025 📍 IIT Ropar

From Chalkboards to Chatbots: Exploring Teacher and Student Readiness for Agentic AI in Indian Schools

Status: Presented (Pending Publication)

Link TBA
IEEE · CINE 2024 📍 KIIT University

Conversational Text Extraction with Large Language Models Using Retrieval-Augmented Systems

Status: Published

Read on arXiv →
IRJMETS Journal 2024

Innovative Fusion of LSTM and Bi-GRU Networks for Enhanced Hate Speech Detection in Social Media

Status: Published

Read Paper →
Springer · ICDCIT 2023 📍 KIIT University

Deep Learning in Industry 4.0: Transforming Manufacturing through Data-Driven Innovation

Status: Published

Read Paper →
Springer · AIIAF 2023 📍 NIT Rourkela

Advancements in Computer Vision and Machine Learning for Food Quality Evaluation: A Comprehensive Review for the Food Industry

Status: Published

Read Paper → Best Poster Award

Articles

2024

Beyond Reality: Virtual Worlds Shaping the Future of Food Science

Food Infotech Magazine — January 2024

Read Article →

Advancements in Food Safety and Quality Evaluation Using Computer Vision and Machine Learning

Food, Marketing and Technology Magazine — March 2024

Read Article →

LLM for Food: from farm to fork

Food, Marketing and Technology Magazine — April 2024

Read Article →

Implementing the Technological Intelligence in Agricultural Produce

Food, Marketing and Technology Magazine — June 2024

Read Article →

Edible and Biodegradable Packaging Innovations

Food, Marketing and Technology Magazine — December 2024

Read Article →

2023

Two Layer Classifiers of MVS for Effective Grading, Safety & Quality Evaluation in Food Industry

Food Infotech Magazine — April 2023

Read Article →

Crunching the number of safer foods: how big data is transforming food safety

Food Infotech Magazine — May 2023

Read Article →

Cracking the Code! Unveiling the Hidden World of Food Safety with MicroWaves & ML

Food Infotech Magazine — June 2023

Read Article →

Revolutionizing Agri-Food Supply Chain: Harnessing the Power of IoT

Food Infotech Magazine — July 2023

Read Article →

Revolutionizing Food Safety: How BlockChain and Lighting Network are changing the game

Food Infotech Magazine — September 2023

Read Article →

Revolutionizing Brain Tumor Detection: The CNN-based Medical Imaging Breakthrough

Informs London Magazine — December 2023

Read Article →

Blog — Writing on Medium

My side hustle and running passion for AI — deep dives, model breakdowns, and industry analysis, published regularly on Medium.

Understanding Latency in Multi Agentic AI Systems

August 15, 2026

Understanding Latency in Multi Agentic AI Systems | DeepDive #1

Model latency is only 30–40% of the total latency shown by the system.

Read on Medium →
Kimi K3: 2.8T most capable model by Moonshot.ai

July 21, 2026

Kimi K3: 2.8T Most Capable Model by Moonshot.ai

Moonshot's flagship 2.8T-parameter model — novel attention and MoE design for efficiency.

Read on Medium →
The Complete Guide to Attention Variants in Transformers

June 9, 2026

The Complete Guide to Attention Variants in Transformers: From Scaled Dot-Product to Flash

A comprehensive overview of transformer attention mechanisms and their evolution.

Read on Towards AI →
The Hidden Engine Behind Fast LLM Inference

May 5, 2026

The Hidden Engine Behind Fast LLM Inference | Why It's Necessary for All Wrapper Applications

Infrastructure optimizations powering rapid LLM responses beyond model speed alone.

Read on Medium →
The Quiet Revolution Happening Inside Your Software

April 19, 2026

The Quiet Revolution Happening Inside Your Software — and Why Most Developers Are Already Behind

How agentic AI systems are quietly reshaping software development.

Read on Medium →
How Perplexity AI makes money on top of GPT, Claude, Gemini

April 11, 2026

How Perplexity AI Makes Money on Top of GPT, Claude, Gemini: Inside the Answer Engine Business

Analyzing a $21B startup's business model built on third-party LLMs.

Read on Medium →
Meta's TRIBE v2

April 2, 2026

Meta's TRIBE v2: The AI That Predicts How Your Brain Reacts to the World

A neuroscience foundation model trained on 500+ hours of fMRI data.

Read on Medium →
Kimi K2.5

February 8, 2026

Kimi K2.5: The Next Milestone in Multimodal and Agentic AI

Moonshot's multimodal model advancing visual understanding and agentic capability.

Read on Medium →
Why LLMs Forget Long Conversations

January 25, 2026

Why LLMs Forget Long Conversations | AI Replacing Humans Narrative

Context window limits keep LLMs from developing genuine memory, unlike human cognition.

Read on Medium →
Compression Is Easy. Preserving Meaning Is Not

January 5, 2026

Compression Is Easy. Preserving Meaning Is Not | Compressed Language Models — Beginner 101

Shrinking the model is the easy part. Keeping it intelligent? That's hard.

Read on Medium →

Authored Books

Available Now ⭐ Top 10 TensorFlow Books 2025 — BookAuthority

Fundamentals of Convolutional Neural Networks with TensorFlow

A deep dive into Computer Vision, Object Detection, AI, CNNs, and TensorFlow.

Computer Vision Object Detection AI CNN TensorFlow

Available on: Amazon, Kindle, Flipkart

Get The Book
Fundamentals of CNNs with TensorFlow
Available Now Amazon KDP · Pothi

Trust in the Age of Agentic AI Economy

A C-suite guide to building trustworthy, reliable agentic AI systems at scale. Drawn from the infrastructure OpenRAG Innovations has built in production — covering hallucination mitigation, retrieval reliability, multi-agent coordination, and the emerging trust frameworks enterprises need as AI systems become autonomous decision-makers.

Agentic AI LLM Infrastructure RAG Systems AI Trust & Safety

Co-authored under OpenRAG Innovations Pvt. Ltd.

Trust in the Age of Agentic AI Economy — book cover