
Imagine opening a small bakery. You would need more than a good cake recipe. You would need ingredients, an oven, a kitchen, someone to serve customers and a way to know whether the business was working.
Building with AI has a similar shape. The model is one important part. Data, computing power, software, people and everyday operations turn its abilities into something useful.
When someone says, “We want to build our own AI,” my first question is: what should it help someone do? That answer determines whether you need to connect an existing model to a product, adapt a model to your needs or train a new one. It also determines what the project should cost.
Let us walk through the whole journey using an imaginary business called MarketMate: an assistant that helps a shop owner answer stock questions and prepare customer orders. The examples and budgets below are teaching illustrations, not claims about a real company.
First, understand what you are building
A model is a mathematical system that learns patterns from examples. A product is the experience someone uses to get a job done. An AI product might contain one model, several models and plenty of ordinary software.
This article focuses on language models, which work with text, and the products built around them. AI also includes models for images, speech, recommendations and forecasting. If MarketMate only needed to predict next week’s bread sales, a forecasting model might be a more suitable starting point than a chatbot.
For a language model, training means adjusting its internal numbers, called parameters or weights. Inference means using the trained model to produce an answer. Training is the learning stage; inference is the working stage. Both need resources, but their workloads and bills differ.
Step 1: Discover the problem and define success
Before choosing technology, I would spend time with the shop owner. What questions arrive each day? Where do mistakes happen? Which tasks take too long? What information is available, and who would pay for an improvement?
Suppose customers repeatedly ask whether an item is available and what it costs. The owner then checks a spreadsheet and types a reply. That gives MarketMate a narrow first job: find the correct item, check live stock and prepare an accurate response.
Write down the current process, the cost of doing it and the cost of getting it wrong. Agree which actions need a person’s approval. For our first version, MarketMate may draft an order, but the owner must confirm it before anything is sent.
The tangible result of discovery is a short product brief: the user, the task, the information required, the limits and the measure of success. “People like our chatbot” is vague. “Owners can prepare correct order drafts with less effort” gives us something to investigate.
Step 2: Choose the right route
Use an existing model
An API is a way for one piece of software to ask another for help. MarketMate could send a request to a hosted model and receive an answer. The provider runs the model’s computers; our team builds the shop experience around it.
This gives us a baseline, meaning a first result against which to compare improvements. We have built an AI application, but have not trained its underlying model. I would describe that honestly when speaking to customers.
Hosted access brings dependencies: provider availability, usage limits, changing prices and rules about how data is handled. Those belong in the product decision.
Give it access to information
Retrieval-augmented generation, usually called RAG, retrieves relevant material and supplies it to a model while it answers. Think of giving someone the right pages of a reference book. In a typical application, this does not change the language model’s weights.
For MarketMate, written delivery policies could be retrieved from approved documents. Live prices and stock should come from the shop’s current records through a controlled software connection. A search result from an old catalogue should never outrank the stock system.
Document search may use embeddings: lists of numbers that help software compare the meaning of passages. Keyword search can also help, especially with exact product codes. Retrieval and training solve different problems and can be combined.
Adapt an existing model
Fine-tuning adds training to a model that has already learned useful patterns. Supervised fine-tuning, or SFT, uses examples of the behaviour we want, such as a customer request paired with a correct order draft. It may help with consistent formatting or specialised tasks.
For MarketMate, the training examples might teach the model to distinguish an item name, quantity and delivery instruction. Changing stock levels would still come from live records.
LoRA is a method that trains small sets of additional parameters while leaving the original weights fixed. It can reduce the resources needed for adaptation compared with updating every weight. It is like fitting a focused adjustment to an existing machine.
An available set of model weights is not automatically permission for every commercial use. Read the actual licence and any conditions on redistribution or derivative models.
Train a model from scratch
Pretraining starts a model with newly initialised weights and teaches it from a large collection of examples. Many text generators learn by predicting the next token, a small unit of text that might be a word, part of a word or punctuation. A transformer uses attention mechanisms to process relationships within its input.
A small educational model and a broadly capable foundation model are very different projects. The latter requires substantial data, experiments, computing capacity and specialist work. Reusing a suitable pretrained model can greatly reduce the training burden.
I would require a clear reason for starting from scratch: an unmet capability, a distinctive data opportunity, a research objective or constraints that existing models cannot satisfy. A desire to put our name on a chatbot would not justify that investment.
Step 3: Build the data foundation
For MarketMate, I would begin with product names, units, stock records, delivery rules and examples of real customer requests that we are authorised to use. Each source needs an owner and a record of where it came from.
Then comes the unglamorous work: removing duplicates, fixing inconsistent units, checking labels and identifying missing information. “Two cartons” means little if nobody has recorded how many items a carton contains.
Keep training examples separate from validation examples, which guide development, and a final test set, which checks the finished candidate. Near-duplicates should not slip between them. Where records repeat over time, split by the relevant customer, document or period so the test resembles a genuinely new situation.
I would also ask whether examples represent the intended users. A shop assistant for Lagos or London may encounter local product names, abbreviations and mixed language. We need relevant examples and reviewers who understand them.
Generated examples can help explore variations, but I would review them against real work. Thousands of polished examples of the wrong behaviour would simply make the wrong lesson more consistent.
The output is a versioned dataset with its origins, permitted uses, review process and splits documented. Keep training data free of unnecessary secrets, credentials and personal details.
Step 4: Define the tests before the expensive work
Evaluation gives the team a shared definition of good performance. Use realistic tasks, clear marking rules and checks for failures with serious consequences. For agents, verify the actual outcome of actions. Repeat important cases and compare changes against the baseline.
Our MarketMate tests would include a misspelt product name, an unavailable item, an ambiguous quantity and a request to view another shop’s orders. An acceptable answer must respect both the facts and the user’s permissions.
I would measure whether the order draft is correct, whether the owner needs to edit it and how long the whole task takes. A fluent response does not compensate for the wrong price.
The tangible output is a test collection and agreed release criteria. These should help us decide whether extra training produces enough value to justify its cost.
Step 5: Assemble the technology and infrastructure
Infrastructure is the collection of machines and services that keeps the work running. Think of the bakery’s kitchen, electricity, storage and delivery arrangements.
Compute and memory
GPUs are processors well suited to doing many numerical calculations at once. CPUs handle general computing work. GPU memory, often called VRAM, holds the numbers and working information needed during model execution.
Memory planning must include more than the stored model. Training may also hold gradients, optimiser state and intermediate results. The exact requirement depends on the training method and precision.
For scale, seven billion parameters stored at two bytes each require about 14 GB for the weights alone, using decimal units. That arithmetic does not mean a 14 GB device can train or comfortably serve the model. Working memory, context length and simultaneous users add overhead.
A small experiment may fit on one machine. Larger work may split data or model computation across several GPUs. Those machines must exchange information quickly; buying more GPUs does not guarantee a matching increase in speed.
Storage, networks and software
MarketMate needs durable storage for datasets and model files, a database for shops and orders, and fast enough access to keep its compute busy. A checkpoint is a saved training state that helps us resume after interruption. Losing checkpoints can turn a temporary failure into a costly restart.
A practical language model training stack can use Python, PyTorch and Hugging Face libraries. PyTorch provides the training machinery; Transformers provides model implementations; Datasets helps organise examples; TRL and PEFT support training and adaptation workflows. These are examples of components, not a requirement to install every tool.
I would also record code versions, configuration, dataset versions and experiment results. Packaging the environment consistently helps another engineer reproduce the work. Capacity management, monitoring and recovery become more important as the number of machines grows.
Cloud rental avoids purchasing hardware immediately. Owning hardware brings purchasing, power, cooling, maintenance and spare capacity decisions. The useful comparison is the total cost of meeting the workload, including the people operating it.
Step 6: Train in small, measurable experiments
The basic training loop is surprisingly straightforward. Give the model examples, let it make predictions, measure the error with a loss function and use an optimiser to adjust the trainable parameters. Repeat. A batch is a group of examples processed together; an epoch is a pass through the training dataset; the learning rate controls the size of updates.
For an adaptation project, I would begin with a small run to check the data format, memory use and whether the model can learn the intended task. Save checkpoints and watch validation performance. A lower training loss alone does not prove that customers will get better results.
If the model improves on practice examples but performs poorly on unfamiliar ones, it may be overfitting: learning the practice paper too closely.
Which training stages are actually necessary?
A project can use continued pretraining to expose an existing model to more domain text, then supervised fine-tuning to teach useful response behaviour. These stages need a demonstrated purpose; not every project needs both. Training a new model from scratch additionally requires choices about the tokenizer, architecture, parameter count and data mixture.
Preference training uses judgements about which response is better. DPO, or Direct Preference Optimisation, learns from preferred and rejected answers to the same request. Other approaches use reward signals, including reinforcement learning methods. Such training needs careful evaluation because a reward can encourage behaviour we did not intend.
For MarketMate, I would try better instructions and data access first, then a small SFT experiment if a repeated behavioural weakness remained. Each extra stage should earn its place through improved results.
Step 7: Work out what it will cost
The full budget has several parts: people, data collection and review, experiments, computing, product development, security, deployment and ongoing support. Model training is one line in that budget.
All figures below are illustrative calculations unless explicitly described as a published price. They exclude tax and currency conversion. The assumed hours are not estimates of how long a particular model will take to train.
Training compute
A simple calculation is: number of GPUs × hours used × price per GPU per hour.
Lambda’s published single-GPU listings show an H100 PCIe at US$3.29 per GPU hour. At that listed rate, one GPU running for an assumed 20 hours costs US$65.80. Four identical 20-hour experiments would cost US$263.20 in instance time, if that configuration and rate remain available. This excludes separate project costs and is not a promise that a chosen training job will fit or finish within 20 hours.
For a larger hypothetical job, 64 GPUs × 720 hours × an assumed US$4 per GPU hour equals US$184,320. The US$4 rate is a teaching assumption, not a supplier quote. Nor does that spend guarantee a useful foundation model.
Estimate run time from a short measured trial using the intended model, data length and hardware. Include failed runs, evaluations and experiments when setting a spending cap.
The monthly inference bill
For a token-priced API, estimate input and output usage separately: input tokens divided by one million times the input rate, plus output tokens divided by one million times the output rate.
Imagine 100,000 requests each month, each using 2,000 input tokens and 500 output tokens. At assumed prices of US$1 per million input tokens and US$4 per million output tokens, the model bill is US$200 plus US$200, or US$400. These rates are invented for the calculation, not advertised model prices.
Long conversations, retrieved documents, retries, tool use and additional model calls can increase usage. Audio, images and other services may use different billing units.
For rented hardware, an instance left running for 720 hours at the published US$3.29 hourly rate would cost US$2,368.80 before other costs. That does not establish whether it could handle the same workload. Serving capacity and response quality must be measured.
The cost of reaching a product
Here is an illustrative small pilot budget in pounds.
Assume 40 engineering days at £500 per day: £20,000. Add 10 days of data and domain review at £300: £3,000. Allow £1,000 for experiments and hosting, and £2,000 for testing and security review. The subtotal is £26,000; a 20 per cent contingency makes £31,200.
Those are planning assumptions, not market rate claims or an Amakora quotation. They describe a deliberately narrow pilot using an existing model. They do not fund a foundation model programme. A founder doing the work personally may spend less cash, but the time still has an economic cost.
For a training programme, build a separate budget around the dataset, model size, measured run time, number of experiments, evaluation and serving requirements. “Training from scratch costs X” is not a useful universal statement.
The business measure I would track is cost per successfully completed task: total running cost divided by successful outcomes. Include human review and support. A cheap answer that needs extensive repair may be expensive work.
Step 8: Turn the model into a usable service
Serving software loads a trained model and handles incoming requests. vLLM is one example; its documented features include continuous batching and memory management for serving. Batching lets a server handle work from several requests together.
Quantisation stores or computes model values at lower precision, which can reduce memory requirements. Compatibility, speed and quality depend on the method and hardware, so the compressed version needs testing.
Distillation is another option: training a smaller model using supervision from a larger one. It adds its own training and evaluation work. For MarketMate, I would consider it only once we knew the task well enough to judge what could be lost.
We must measure response delay, requests handled per second and performance during busy periods. Plan for capacity limits, timeouts and recovery. A backup model is only useful if it also meets the required standards.
Step 9: Build the product around it
Now we give MarketMate a usable home. The shop owner signs in, connects the catalogue and sees a simple screen for customer requests. Behind that screen, our application checks which shop the owner belongs to and retrieves only authorised information.
The model helps interpret the message. Ordinary software validates product identifiers and calculates totals. The owner reviews the proposed order before it is submitted. The application records whether submission actually succeeded and avoids creating a duplicate if a request is retried.
This is where design, integrations and operational judgement become visible. The product needs clear error messages, a way to correct mistakes, onboarding, support and, if sold commercially, billing.
I would enforce access permissions in software outside the model. Instructions such as “never reveal another shop’s data” are useful guidance, but they should not be the only barrier protecting it. Treat uploaded documents and incoming messages as untrusted content; they must not be allowed to grant themselves authority.
Make deliberate decisions about data retention, deletion, encryption, secrets and supplier access. Check the applicable data and model terms before launch. Collecting every conversation forever would be a poor default.
Step 10: Pilot, release and keep learning
I would pilot MarketMate with a small group of shop owners and a clearly bounded task. Watch them use it. Record corrections, failures, time saved and whether they return to it. A polished demonstration can conceal a confusing daily workflow.
Before expanding, agree who owns incidents, how to disable unsafe actions and how to return to the previous working version. Monitor the full service, including integrations, model behaviour and costs. Add new failure cases to the evaluation collection.
If releasing adapted weights, package the model or adapter with its required base version, tokenizer, configuration and usage instructions. A model card should explain intended uses, training information, evaluation results and limitations.
For the product, the release package also needs the running application, tested connections, support arrangements and an operating budget. Customer feedback should enter a reviewed improvement process; it should not silently become training data.
Who needs to do the work?
The work spans several skills. Someone must understand customers and the domain. Someone must organise data and evaluate quality. Model training needs machine learning expertise. A functioning product needs application engineering and design. Reliable deployment needs infrastructure and security ownership.
A small team can combine responsibilities, and specialists can contribute for specific stages. The important thing is that each responsibility has an owner. Training a broadly capable model from scratch creates a much larger research and operations commitment than a narrow API-based product.
What should exist at the end?
For our imaginary MarketMate, a tangible first product would let an authorised shop owner ask a stock question, receive a grounded answer and approve a correct order draft. It would have a working interface, current data connections, measured quality, clear limits and someone responsible when it failed.
Along the way, discovery produces the product brief. Data work produces reviewed datasets and access rules. Experiments produce a measured baseline and, where justified, a trained model. Engineering produces the service and interface. The pilot produces evidence about usefulness and cost.
My view is that building AI requires the discipline to connect each technical choice to the person who will depend on it. Model ownership can be valuable, but its value must show up in better capability, control, economics or customer outcomes.
Start with a problem worth solving. Choose the model strategy that serves it. Carry the work all the way through to a product someone can use with confidence.
The prompt that turns an AI idea into a first brief
If you have an idea but are unsure where to begin, use this prompt to organise the questions:
I want to help [specific user] do [specific task]. Today they use [current process]. The information available is [data], and my constraints are [budget, time, privacy and skills]. Help me compare a simple software solution, an existing model with tools or retrieval, fine-tuning and training from scratch. Explain which option to test first and why. List the assumptions I must validate with users. Propose a small pilot, the tests it must pass, its required infrastructure and a budget formula. Separate verified facts from assumptions. Do not invent prices or promise results.
Treat the response as a draft for investigation. Speak to users, inspect the data and check supplier prices before committing money. The prompt is useful because it exposes decisions; it cannot make those decisions responsibly without evidence.
Your first week of building
On Monday, speak to potential users and choose one recurring problem. On Tuesday, map the current workflow and check whether the necessary data is available. On Wednesday, write realistic test cases and define an acceptable result. On Thursday, try the simplest baseline. On Friday, compare it with the current process and decide whether to continue, change direction or stop.
This is a suggested discovery rhythm, not a promise to train and launch a model in five days. A useful first week should leave you with a clearer problem, evidence about feasibility and a sensible next investment.
Reflection for the week
There is something deeply satisfying about building a system that another person can use to make their day easier. The responsibility is to follow the idea beyond the exciting demonstration and into the ordinary details: the wrong input, the monthly bill, the person waiting for help.
Ambition becomes more credible when we can explain the work, the cost and the evidence behind it. That applies whether we use an existing model or train one ourselves.
Want support moving from an idea to a working system?
Amakora’s System Architect Fellowship is relevant for readers who want practical experience with systems, AI infrastructure, automation and technical delivery. The AI Builder Bootcamp provides a pathway into workflows, prototypes and AI-enabled systems.
If you lead product work, the Technical Product Management Fellowship connects APIs, system design and AI product thinking. Choose the pathway that matches the responsibility you want to develop, then discuss the specific project and depth of training you need.
Explore Amakora’s learning programmes or apply for a programme. Organisations with an implementation brief can also talk to Amakora about architecture, integrations and a focused pilot.
Until next week,
Tochii
Founder, Learn with Tochii
Inspire · Educate · Empower
contact@tochukwuachebe.com
tochukwuachebe.com
Sources and further reading