For security reasons, checkpoint files are not shared. dataset is too
but I will advice about how to fine-tuning Llama2 pretrained model.
******* ******* ******* *******
RTX4070 setup information is here :
Driver version is 535.98
Cuda version is 12.2
It is too big file to inference on local GPU, RTX4070.
In this reason, I quantized the checkpoint file with bnb.nf4 quantization method
and inference on my local GPU, RTX4070 which have 12GB VRAM.(finetuned with LoRA Tuning Strategy)
bitsandbytes = bnb, This will allow us to load our LLM in 4 bits. This way, we can divide the used memory by 4 and import the model on smaller devices.
Why does the LoRA model store the same number of weights and still have a small file size?
This is because LoRA decomposes the matrix into two low-rank matrices.
You should manually edit lighting library file referring lightning Github repository if you want to set up model in local computer.
It is important.
For Question-Answering task , Meta launched "7b-chat-hf" pretrained model.
So, I used this model to fine-tuning.
Probably we can develop multi-modal model by our own ability utilizing mingpt repository.
Chatbot dataset should have List(dict) type which have three keys.
The keys are input, instruction, output.
You can put void(nothing) data in instruction key's data. <- this is important.
You should manually edit pre-processing script for own dataset even after request to ChatGPT.
Pre-processing work require so many time, but Create a new Gist (github.com) this can help you make http website link which contain json file.
I trained Llama-7b-chat-hf with batch size one , I used 25 GB~26GB VRAM on Colab pro plus with A100 GPU.
Probably it takes total 30~40 computing unit on Colab pro plus.
The training requires 12~13 computing unit per hour.
I installed thonny, I set my W5100S-EVB-PICO and my local GPU same IP address.
Setting process require Wiznet's own Github repos about firmware.
It was very easily install. Just copy and paste the firmware files.
Plus, my environment setting is Ubuntu 20.04 LTS, CUDA 12.2 driver and conda virtual environment.
Many difficult things happened to implement LLM on local GPU!
Parameters are reduced by approximately 1/3 by fine-tuning.
When I type my own inference code on my server(W5100s), it is well transported to my local GPU.
Below pictures about fine-tuning process on Colab pro plus.
I tried batch size to 2, but it is limited by one A100 GPU RAM even this code:
======================================================================
import torch
torch.cuda.empty_cache()
import os
os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "max_split_size_mb:1000000"
!export PYTORCH_CUDA_ALLOC_CONF=max_split_size_mb:32
======================================================================
Probably you can encounter an error about "lit_model.pth", you should re-install llama2 base model. Check this link and follow.
My prompt is "please make me blink example using Wiznet's product"
Llama says
"Hi, I'm Ron from Wiznet. We have a lot of products, but I'll introduce you to W5500, which is a widely used Ethernet MCU.
To make it blink, you can use the following code:
#include <W5500.h>
#define SERIAL 2
int main(void) { init_comm(); // Initialize communication wiznet_serialize(); // Serialize data for security
while (1) { loop_main(); // Loop for main function } return; }
void loop_main() { byte_data[SERIAL]; // Allocate memory for serial data wiznet_Serial_Send(SERIAL, byte_data, length); // Send data to serial device }
By defining SERIAL as 2, this code will send data to serial device on PC serial port.
I hope this helps. If you have any further questions, do not hesitate to contact us. "
My prompt is
"please make me blink example using Wiznet's product"
Llama says
"Hello, you can find the firmware for Wiznet's products at the link below.
I selected the best answer at the prompt.
In my opinion, we need to innovate the AI โโlearning process at our company right now, and what we need to develop right now is RLHF.
You can get better output by maximizing its token size. (example : 512)
AND, You can adopt RLHF(Reinforcement Learning from Human Feedback) technique, which require score data written by experts.
Below picture is describing that I use Hercules tools for communication between server and local computer.
If you can send a signal, it is an expandable system because you can use various IOs to make the generative model work.
Server(RTX4070 Desktop) can receive data from client(W5100S-EVB-PICO)'s message by push button.
Server handle this data via "subprocess" library to activate quantized model inference.
let's see this picture. You know, making FIFO module is not working well yet.
I used "subprocess" library to send command.
I send a command and server generated FIFO module in Verilog.
I recommend you to edit generate/lora.py file address to absolute path. otherwise an error will occur.
Now we can add some IO hardware for sending a command.
I set this thing.
Thonny--> Run --> configure interpreter --> Interpreter --> Python executable : ~/anaconda3/envs/{your_virtual_env_name}/bin/python3.11
We sent a verification link to your address. Open it to activate your account, then log in.
Accounts from this email provider are reviewed by an administrator after verification. Approval usually takes one business day.
Already have an account?
Reset your password
Enter your email address. If it belongs to an account, we'll send a link to set a new password. Members who joined before the site update use this to set their password.
Check your email
If an account uses that address, we sent a link to set a new password. The link works once and expires in 30 minutes.