EXCEEDS logo
Exceeds
shaharmor98

PROFILE

Shaharmor98

Worked on integrating LoRA and PEFT adaptation into the nv-auto-deploy/TensorRT-LLM repository, enabling adapter-based model experimentation and deployment. Developed an end-to-end LoRA flow that supports execution of LoRA adapters from model loading through inference, updating both C++ bindings and Python configuration to handle LoRA parameters. Enhanced the PyExecutor and TensorRT-LLM components by incorporating PEFT caching and modular architecture changes, allowing for flexible, adapter-ready deployments. Leveraged skills in C++, Python, and deep learning to improve cross-language integration and resource management, ultimately accelerating inference deployment workflows and supporting more efficient experimentation with parameter-efficient fine-tuning techniques in production environments.

Overall Statistics

Feature vs Bugs

100%Features

Repository Contributions

2Total
Bugs
0
Commits
2
Features
1
Lines of code
857
Activity Months1

Your Network

1814 people

Work History

April 2025

2 Commits • 1 Features

Apr 1, 2025

April 2025: Focused on enabling LoRA/PEFT adaptation across PyExecutor and TensorRT-LLM to accelerate experimentation and deployment of adapter-based models. Delivered end-to-end LoRA flow, enhanced PEFT caching, and updated core components (C++ bindings and Python config) to support LoRA parameters, driving faster time-to-value for inference deployments and greater modeling flexibility.

Activity

Loading activity data...

Quality Metrics

Correctness90.0%
Maintainability80.0%
Architecture90.0%
Performance80.0%
AI Usage20.0%

Skills & Technologies

Programming Languages

C++Python

Technical Skills

C++Deep LearningExecutorFull Stack DevelopmentLoRAModel OptimizationPEFTPybindPythonResource ManagementTensorRT

Repositories Contributed To

1 repo

Overview of all repositories you've contributed to across your timeline

nv-auto-deploy/TensorRT-LLM

Apr 2025 Apr 2025
1 Month active

Languages Used

C++Python

Technical Skills

C++Deep LearningExecutorFull Stack DevelopmentLoRAModel Optimization