Adding JSON mode to any model and that too without prompts

Learning about how to add JSON mode to any model and dont just solely on prompts

October 10, 2025 · 5 min · Mohit Dulani

Learnings from astro app AS

Load balancer Building up load balancer Get the input metrics sorted : Latency Cost Uptime Fit a linear equation that matches your score function : SF = W_1 * latency + W_2 * cost + W_3 * uptime where W_1 + W_2 + W_3 = 1 AB testing Create 2 buckets 80/20, whenever a user comes check if its in A bucket or B bucket or assign it one based on the rule so now it becomes a cohort to choose from and see which is a primary and which is a secondary users and experiment with them accordingly !! ...

September 10, 2025 · 4 min · Mohit Dulani

Deep learning Optimizers

Gradients visualised Second order derivative Maths behind Single variable Things are very simple in single variable Function definition : f(x) = x^2 First derivative : f'(x) = 2x Second derivative : f''(x) = 2 To find the Minima in single variable the f’(x) = 0 , and f’’(x) >= 0 , these 2 conditions are enough to find minima Multivariable /Multivariate 2x + 3y = 10 (multivariable eq) x2 + y2 = 16 (multivariable eq) ...

September 9, 2025 · 11 min · Mohit Dulani

Virtual Machine

Connecting to a Azure VM using ssh and RDP Configure a VM based on the specs Connecting with SSH Directly connect using ssh create an inbound rule for ssh, you get the ssh key ! Connecting with RDP RDP , install ubuntu-desktop, 3389 is the port for RDP , 22 for ssh , 80 for http , 443 for https ! Azure VM RDP Setup Guide 1. Install Desktop Environment (Ubuntu Desktop) Install desktop environment (if not already installed): sudo apt update sudo apt install ubuntu-desktop 2. Install and Configure XRDP Install xrdp: ...

September 3, 2025 · 2 min · Mohit Dulani

Home Lab

Kubernetes Kubernetes is a manager of pods, those pods could be running on different machines and we can have multiple roles to this, one could be control panel other could be worker role So k8s takes your machine ( or you can define via virtualisation, the amount / ratio it should take) and then it orchestrates it, so you have to define a .yaml file and k8s will take care of on which cluster to run this and how many pods to spin up for this , everything is taken care of just you need namespace, entry file / project etc all this written in a yaml file. ...

August 31, 2025 · 7 min · Mohit Dulani

Post training methods in LLM using RL

Tags : PPO RLHF Maths Reinforcement learning, here the agent takes / decides some action to take based on the current state and other variables present at timestep t, and then its takes that action and a reward is followed and weights are updated based on the rewards received by model Consider this basic hello world example of RL State : Any place / position where the agent can be Action : Up , down , left , right these are the action the agent can take ...

August 23, 2025 · 6 min · Mohit Dulani

LLM's loss curve and compute

Small LM vs Large LM Small no. of parameter model has to choose between what knowledge to keep in parameter space and what to ignore and due to small no. of params the model tends to ignore most of the ood knowledge and keep the one that is commonly occuring ( small fields are ignored / tail knowledge is ignored ) Multi-task learning : grammar maths punctuation .. etc Heuristics : Small LM unwilling restricts itself to a smaller set of tasks , it can only improve a particular set / learn a particular subsets ( grammar , punctuation only ) while a large language model can learn both about tail knowledge / tasks like maths, punctuation and world knowledge ...

August 16, 2025 · 2 min · Mohit Dulani

Django

Building a scalable backend arch Using Django / django We start with a django Project, Using : $ django-admin startproject <project-1> <dir_name> A single django project contains its own views, urls, asgi , wsgi , manage.py files and a project is often called as service so these 2 things projects and services are same only, and the microservices architecture is the one where each service aka project can be scaled up independently .. ...

July 20, 2025 · 2 min · Mohit Dulani

Learning Go-lang

Learning GO to get the minimum latency and get things right from start .. Go docs Go Tour The whole file needs to become a package and to make it whole a package we wrap that up in package main and this makes the whole code as a single package that is then converted to binary We create package to help us import those in other files and the importing helps in code seperation, else we would have to write the whole code in a single main.go file.. ...

July 18, 2025 · 9 min · Mohit Dulani

Low Rank Adaptation

LoRA This is beneficial only if the rank, r « d, where d is the dimension of the matrix Rank of a matrix No. of Linearly independent rows/ columns are ranks A matrix of size 4 x 5 with rank = 2 , can be broken down to ( 4 x 2 ) x ( 2 x 5) reducing the total no. of params from 20 to 18 And when applied at a large scale for sizes 1024 x 1024 or 2048 x 2048 , and rank = 2 or 4 the size decreases exponentially from (1024 x 1024) to (1024 x 4 x 2) by 128x times ...

June 7, 2025 · 2 min · Mohit Dulani

SmolVLA paper from HG

SmolVLA Vision language action paper used for training real world Robots for task Input : Images of surrronding, task explanation in text, state ( senserimotor snapshot at a given time ) Sensorimotor states are projected into a single token using a linear layer to align with the token dimension of the language model. Ex: i } n p u t " " " ] s i i s ) m n t = a s a g t t 0 0 0 0 0 # { e r e . . . . . " u " 1 0 8 4 0 : c : 2 0 5 5 , t , , , , i i n 0 m o p - 0 0 . a a n . 1 . . 0 n g " a . 0 1 , y e : r 0 1 2 _ r 5 , , 0 o t " a , . t e P y 0 0 0 h n u ( 0 . . , e s t [ . 0 3 r o 4 0 3 1 r t 4 , , . r , h , 0 e e 0 , l 0 . e # r . 0 v e 3 0 a e d 1 , n . , t g c 0 . u - . s , b 0 0 e e . 0 n [ 2 , s 3 i 2 o , n , 0 r . 2 t 0 0 r 2 h . 0 e 4 e 0 , a , 8 d b , 0 i 2 i . n 2 n 1 0 g 4 " . 0 s ] , 5 , 7 R , G B i # # m # # # a j j g o o g e e e i i r n n n n i d d f t t p - - r p e e o p v e f f m o e r f f s l e e r i o o c c o t c p t t b i i e o o o o t n r r t n i ' s e ( p o s s 1 o r ( . s i c 7 ( 0 i e a ) 7 = t n m ) o i t e p o a r e n t a n i , ( o x n 0 , . ( 0 y q = , u c a l z t o ) e s r e n d i ) o n ) Output from Action expert that predict what action it should take next: Flow matching is a way to train the action expert so it can generate smooth, realistic action sequences quickly and efficiently. ...

June 7, 2025 · 4 min · Mohit Dulani

Thinking and fiddling out with random ideas

Learning new ideas and fiddling with them When you visit a site what happens ? *The html , js , css gets transferred to the browser using http request and then the chrome engine runs that and display the content How does image upscaling works? How does real time stt works ? We store the 500ms chunk in /tmp file , send it to transcribe and then repeat this until we close it .. or a VAD (voice activity detection) is used to check till when the voice is detected .. ...

May 11, 2025 · 2 min · Mohit Dulani