
Aru Sharma
B.E. Information Technology Student at
University Institute of Engineering and Technology, PU
About Me: I am passionate about building intelligent systems that bridge the gap between human communication and machine understanding. I build multimodal AI systems, contribute to open-source projects, and explore the intersection of NLP and Computer Vision. My approach to engineering is to study where current solutions fall short in real-world applications, and develop practical improvements that make AI more accessible and useful.
Experience: My journey involves significant contributions to open-source ecosystems. I worked as a summer intern at Google Summer of Code with Mifos Initiative, developing multi-agent bots. I interned at Summer of Bitcoin and contributed to Bitcoin Transcripts, and was an LFX Mentee at CNCF WasmEdge. I contributed to DocETL at UC Berkeley's EPIC Lab. I have also worked as an AI Engineer with Bennett Legal and HomeHive AI, and conducted risk analysis for crypto tokens.
Community: I lead the OSS club (Pclub) at my college to promote OSS. I also hosted events like Software Freedom Day, OSS hackathons like FOSSHACK and MOSS HACK in collaboration with Moss (YC F25), and started AISOC so that students can get familiar with how to start contributing to OSS. I am also returning as a GSoC mentor for Mifos, and am mentoring Genesis-KB at Summer of Bitcoin.
Achievements: Selected for the first edition of ESOC'25 under the Open-Source AI for Drug Discovery project. Ranked 15 globally on the NTIRE Image Dehazing and Denoising challenge at CVPR 2024. Published research on Speech Emotion Recognition accepted at the 16th ICCCNT 2025. Published a patent on healthcare and AI using computer vision models.
About
Education
B.E. Information Technology
University Institute of Engineering and Technology, PU
Mathematics and Computer Science
Little Scholars, Kashipur (CBSE)
Current Focus
Interested in long horizon reasoning and self-evolving agents that learn and adapt over time.
Exploring mechanistic interpretability and AI for science to understand how intelligent systems work and apply them to real-world research problems.
Experience
- Worked on building memory infrastructure for long horizon reasoning and better personalisation.
- Developed Tetrix CLI- a tool to review architecture, and security issues and enforce code quality for your project.
- Worked on testing and deploying SOTA Vision algorithms for classification, segmentation and pose detection
- Deployed OSS text to video generation models for in-house testing and benchmarking against Veo3
- Developed a multi-agent bot letting users know the status of Jira tickets, questions related to Slack discussions.
- Developed a full-stack web application using FastAPI, NextJs and Firestore as database and Auth client.
- Designed and prototyped AI-assisted coding tools for Bitcoin using small language models and domain-specific Retrieval-Augmented Generation (RAG).
- Developed data pipelines to ingest knowledge from Bitcoin developer calls, YouTube talks, IRC logs, mailing lists, and forums.
- Contributed User Defined Functions, LLM based data parsing and OCR modules to enhance the usability capability of DocETL.
- Added structured generation support for Open-Source model based backend using Outlines.
- Developed a RAG based chatbot for code assistance using opensource LLMs with Wasmedge runtime.
- Created a pipeline to ingest data from Github repository, augmented it using QnA pairs, summary and then embed this into a Qdrant vector database.
Projects
AI-native open-source analyser for your coding patterns.
Key Features:
- Built a lightweight toolkit to analyse coding sessions from Claude Code using LLMs and SQLite
- Keeps private data on the local machine and creates a profile suggesting strengths, areas of growth, and overall narrative
Technologies:
Knowledge base and Bitcoin education platform for Bitcoin development.
Key Features:
- Transcription engine for processing Bitcoin-related audio and video content
- Web frontend for browsing and searching Bitcoin transcripts
- LLM-based explainer for BIPs and BOLTs
Technologies:
Building a personalised agent that can reason over long term to remember and recall information from past interactions
Key Features:
- Implementation of the EverMemOS paper from first principles
- Keyword as well as semantic based retrieval system combined with reranking mechanism
Technologies:
Implemented a multimodal emotion recognition system using late and gated fusion techniques on audio and video embeddings to classify emotional states.
Key Features:
- Whisper-large-v3 for audio feature extraction
- V-JEPA for video visual embedding extraction
- Gated Fusion Network for combining modalities
Technologies:
Interaccionismo
Shared thoughts on AI, open-source, and building systems
Making Your First Open Source Contribution: A Step-by-Step Guide
Open Source
Towards Personalized Reasoning: Building Agents That Remember
Memory Systems
How to make your First Pull Request to Open Source Codebase
Open Source
How to Write a GSoC Proposal That Almost Gets Accepted
Open Source
You Don't Need a Fancy Memory Layer
Memory Systems
A Deep Dive into EBM, JEPA, and World Models
World Models
From Threads to Tensor Cores: Understanding How GPUs Really Work
Systems
Contact
Get in Touch
Open to full-time positions, and collaborations in AI/ML.