Ask a language model how many letters are in the word “strawberry.” Chances are good it will get it wrong. It might say ten. It might guess eleven. It will sound completely sure of itself, and it will still be incorrect. Ask the same model to calculate compound interest on a loan, or to add two large numbers, or to count business days between two dates, and you will often get a plausible-looking answer with digits that are slightly, dangerously off.
This happens because large language models do not reason about numbers the way humans do. They predict tokens. A token might be a whole word, part of a word, or a single digit. When the model sees “strawberry,” it does not see eight individual letters lined up in a row. It sees a handful of chunks. It has never been taught to count characters, only to predict which chunk of text comes next. The same limitation applies to arithmetic. The model has no internal calculator. It lacks carry logic. It has no real understanding of place value. When it multiplies 148 by 279, it is not performing multiplication. It is pattern-matching against similar expressions it saw during training, guessing what sequence of digits should follow. For tiny sums the pattern is strong enough to work. For anything with real precision, the guess eventually breaks.
Two Jobs, One Bot
Standard prompting methods ask a single system to do two very different things at once. First, understand the logic of the problem. Second, execute the exact math. The model is genuinely impressive at the first task. It can read a word problem, extract variables, map relationships, and plan a solution path. But then it has to serve as its own calculator. That is where the chain frays. A single slipped digit in step three infects every step after it. The logic itself might be perfect, yet the final answer is garbage because the model added wrong.
Program-Aided Language Models, or PAL, solve this by splitting the work. Instead of asking the model for an answer, you ask it for a program.
Here is how the flow actually works. You present the problem. The model figures out the logic, defines the variables, and structures the algorithm. Then, instead of computing the result itself, it writes a short script, usually in Python. That script gets handed off to a real code interpreter. The interpreter runs the logic and returns the exact, deterministic result. The model describes the math. Python does the math.
Executable Reasoning in Practice
Think of PAL as executable reasoning. If a script can solve a problem, let the model write the script.
Consider a concrete example. You need to calculate the maturity amount on a fixed deposit of ₹50,000 at an annual interest rate of 8.5 percent, compounded quarterly, held for seven years. Ask a language model directly, and it might write out a formula, substitute the values, and compute the result in a chain of thought. Look closely, though, and you might find it mishandled the quarterly compounding by dividing the rate incorrectly, or it rounded an intermediate step and carried the error forward. The answer looks reasonable but is off by hundreds of rupees.
With PAL, the interaction changes. You instruct the model to generate Python code that defines principal = 50000, rate = 0.085, time = 7, and n = 4, then computes amount = principal * (1 + rate/n) ** (n * time). The model emits the code. A Python runtime executes it. You get the precise figure, down to the last decimal, every single time. There is no guesswork in the multiplication, no hallucinated remainder, no confident rounding error.
This same pattern applies to date math. Ask a model which date falls exactly 120 business days from today, excluding weekends. A text-only model might count forward and slip on a Saturday. A PAL approach has the model write a script using datetime and calendar logic, then let the interpreter iterate exactly. Data manipulation works the same way. If you need to parse a messy CSV, filter nested JSON, or run a quick statistical transform, the model should draft the logic while the interpreter handles the iteration.
Why This Actually Matters
The shift from prose answers to executable code delivers three practical advantages.
Determinizm. Model językowy, któremu zada się to samo pytanie dwa razy, może zmienić sformułowanie lub cyfrę. Interpreter za każdym razem zwraca ten sam wynik dla tych samych danych wejściowych. Ta stabilność ma ogromne znaczenie w księgowości, logistyce, planowaniu oraz wszelkich obliczeniach inżynieryjnych, gdzie spójność nie jest opcjonalna.
Weryfikowalność. Gdy model podaje trzy akapity rozumowania, musisz przeczytać każde zdanie, aby wyłapać jedną błędną liczbę. Gdy podaje dziesięciolinijkowy skrypt, możesz przejrzeć kod. Możesz zweryfikować, czy wzór na procent składany jest poprawny, zanim interpreter w ogóle go uruchomi. Możesz sprawdzić nazwy zmiennych, wyłapać błędy typu off-by-one i nawet kontrolować wersję rozwiązania. Powierzchnia potencjalnych ukrytych błędów drastycznie się zmniejsza.
Niezawodność. Model trzyma się swojej roli. Robi to, do czego został stworzony: wnioskuje na temat struktury, semantyki i dekompozycji problemu. Maszyna robi to, do czego została stworzona: wykonuje obliczenia z precyzją. Ta separacja odpowiedzialności to dokładnie sposób, w jaki projektuje się niezawodne oprogramowanie. Kompozycja wygrywa z monolitycznym projektem.
Uruchamiaj to jak niezaufany kod
Należy zachować ostrożność. Wygenerowany kod należy traktować jako niezaufane dane wejściowe. Model może napisać skrypt z nieskończoną pętlą, niepotrzebnym żądaniem sieciowym lub operacją na systemie plików, o którą nie prosiłeś. Zawsze uruchamiaj te programy w odizolowanej piaskownicy. Używaj kontenerów z ograniczonymi uprawnieniami, funkcji serverless bez dostępu do sieci lub ściśle kontrolowanych środowisk z ograniczonym czasem procesora i bez trwałej pamięci masowej. Bezpieczeństwo nie jest tu tylko przypisem. Jest częścią projektu systemu.
Gdzie PAL błyszczy, a gdzie się kończy
PAL świetnie radzi sobie z matematyką, datami i manipulacją ustrukturyzowanymi danymi. Eliminuje błędy mechaniczne, które nękają rozumowanie oparte wyłącznie na tekście.
Nie naprawia jednak błędnej logiki. Jeśli model wybierze niewłaściwy wzór,
