Frontiers for Young Minds Grades 9 and up🔬 Science💡 Technology

How Do We Cook Up Artificial Intelligence With Optimization Algorithms?

Abstract

Have you ever followed a recipe to make your favorite cake? Sometimes it turns out perfect, and sometimes you must adjust the amount of sugar or baking time to improve it. But what would you do if you needed to develop a brand-new recipe from scratch? Building artificial intelligence (AI) is similar, but requires a much more complex recipe! To “cook up” smart AI, scientists combine billions of instructions in just the right way, so the program can recognize patterns, solve problems, and reason like a human. However, it is nearly impossible to guess the best sequence of instructions in advance. Optimization algorithms are like helpful chefs in the kitchen, testing how well the AI performs, figuring out what is wrong, and tweaking the recipe to improve it. Optimization algorithms help computer scientists discover better versions of the recipe, until the AI becomes smarter and more powerful—just like perfecting a delicious cake!

Cooking Up Intelligence

We all like to eat. But most of us do not just survive by eating what we can find in nature, straight from the bush (or bone). Instead, we break up the foods we collect into different parts, which we prepare, combine, and cook to produce a tasty and nutritious meal. Many things people like to eat—foods like salad, cake, chicken soup or a Sunday roast—consist of many ingredients and cooking steps, prepared just so. Over time, people have learned recipes: the best ways to prepare and combine ingredients to make tasty and nutritious food.

Artificial intelligence (AI) works in much the same way: some data are collected, then split up into little pieces. The pieces are processed and recombined by the AI, until finally the AI gives us what we want—it tells us what is in a photo, or corrects an essay, or even makes a robotic arm pick a flower. So, congratulations! You have won a ticket to take a special tour inside Willy Wonka’s mysterious AI factory. In this article, we will peek inside the secret “kitchen” where AI is made.

What exactly is AI? It is a kind of computer program that acts on data. In other words, this type of program takes some kind of inputs—for instance, some text, a picture, a live video feed, or maybe a DNA sequence—and transforms them into outputs: a response to the text, a description of the picture, a count of the people in the video, or a prediction about what type of protein the DNA codes for. This type of program is frequently called a model, because that is what it does: it models the relationship between what goes in and what should come out. The models that make up most of what we think of as AI today are a special kind of model called an artificial neural network, or simply neural network (Figure 1).

Diagram showing a process where a cat image is converted into binary data, which serves as input to a multi-layered neural network. The output identifies “cat” as the correct classification among dog, bird, and turtle options.
  • Figure 1 - A neural network is a set of interconnected neurons arranged in layers that work together to model complex patterns in data.
  • In this picture, the input (a photo of a cat) is passed to the first layer of neurons (blue circles), each of which combines the inputs and generates a new value, the output. This is then passed to each neuron the next layer, until finally the output, “CAT” is calculated.

For the AI to work, first, the inputs are digitized, that is, converted into a bunch of numbers that the program understands very well—for instance, the color values of all the pixels in a picture or video. Then, these digitized inputs are passed through the neural network. Most neural networks are made up of layers, which in turn are made up of neurons: computational units that take inputs, multiply them by special numbers called parameters that are different for each neuron and then add the results together. If this sum is above 0, the neuron passes it to the next layer of neurons and so forth, until we get the final answer. The crucial thing is that each neuron has its own parameters, meaning it has its own job in the network. For instance, in a neural network trained for identifying images in pictures, one neuron might be looking for shapes that look like eyes, while another one could be looking for car wheels (Figure 2).

Diagram illustrating a computational process where three colored inputs enter a blue circle labeled "Parameters and Computation." Inside, three parameters are shown. Computation involves a weighted sum and a filter that outputs a value only if the sum is greater than zero, otherwise outputs zero. Three outputs exit to the right.
  • Figure 2 - A neuron is a computational unit that does two things.
  • First, it combines all of its inputs by multiplying each by its corresponding parameter and then adding up the results. Second, it checks if the sum is positive, and if it is, sets the output to the sum—otherwise, it sets the output to 0.

The Patient Chef Inside AI

But where do each neuron’s parameters come from? Picking the parameter values for each neuron is a bit like designing the optimal recipe for something you want to make. Even a simple recipe, such as the perfect mix of water, lemon juice, and sugar for lemonade, could take several tries to find the right combinations. As you can imagine, the more ingredients and combinations you have, the trickier it becomes to balance everything just right.

Now, imagine you are trying to bake a cake with a thousand different layers, frostings, and fillings, all of which must have the perfect flavor and texture so that each bite is perfect and the cake does not collapse on itself! That is what modern AI models are like. These models, which can recognize faces, write stories, or drive cars, have billions of parameters that all affect each other in complicated ways. Because the networks consist of neurons that feed into each other, changing one parameter can subtly change the effect of hundreds of others. It is simply impossible to guess the perfect parameter values in advance [1].

That is exactly where optimization algorithms come in: our tireless, patient chefs who know how to taste each version of the recipe, measure approximately how far it is from perfect, and adjust the recipe in a smart, systematic way to make the next try a little better than before.

Once you taste the result of your current recipe, you can usually tell when something simple went wrong: maybe it is too salty, too sweet, or overbaked. You can connect those problems directly to what you did during cooking. But not every mistake is that easy to spot. Sometimes the final taste is just off, and you have no idea which of the many steps went wrong—or perhaps several things did. It gets even trickier because ingredients can go through many transformations. For example, the flavor of chocolate changes depending on temperature and mixing, and it is hard to tell exactly which step caused the final taste to change. In AI, this is the difficult part: changing the instructions in the early stages of the recipe might affect the final result in ways we cannot easily predict. That is why we rely on optimization algorithms. After every “baking attempt,” the algorithms measure how well the AI performs, then gently tweak each instruction step to improve the outcome.

But how do optimization algorithms know how to adjust the parameters? That is where mathematics comes in. In technical terms, optimization algorithms make these tweaks using something called gradient information [2, 3]. In a nutshell, the gradient is the mathematical answer to the question, “if you could make a tweak to every parameter to make the result a little better, how much should you change each one?” Gradient information tells the optimization algorithm how to adjust the parameters. In many cases, by repeatedly calculating the gradient and making the adjustments, the result improves until the model gets very close to one of the best possible outcomes—like climbing a hill one step at a time until you get to the top. In the cooking analogy, the gradient is like a detailed list of tasting notes that tells you how to make the cake just a tiny bit better: a little more sugar in the frosting, a bit less flour in the biscuit, and turn the oven down slightly. For neural networks, this becomes “make this parameter just a bit bigger and make that parameter just a bit smaller.” If the gradient is large, this is like an obvious mistake you could taste right away. By contrast, smaller gradients are like subtle flavor differences that only a careful chef would notice (Figure 3).

Diagram illustrates an analogy between baking and artificial intelligence, showing inputs as baking ingredients, processed through a neuralnetwork model, producing cake as output, with evaluation and feedback loops guiding recipe adjustments based on tasting notes.
  • Figure 3 - To train a neural network, scientists repeat the following steps until the model reaches the desired performance.
  • First, they pass the input features through the AI model to produce an output. Next, they evaluate this output to measure approximately how far it is from the optimal result. Based on this evaluation, they compute the gradient information that tells them how to adjust the parameters of each neuron to improve the model. After updating the parameters, they run the process again with the modified values, repeating these steps until the performance is satisfactory.

Overall, the training process of a neural network follows a sequence of steps that are repeated many times to gradually improve the model:

• Pass the input features through the AI model to produce an output.

• Evaluate the output to measure how far it is from the desired result.

• Compute the gradient information based on this evaluation.

• Adjust the parameters of each neuron to improve the model.

• Repeat steps 1–4 with the updated parameters.

Thanks to optimization algorithms, we can keep refining our AI “recipes” until they become smarter, more accurate, and more powerful, just like perfecting a delicious cake through patience and practice.

Is The Recipe Ready for Guests?

Now that you understand how to optimize the parameters of an AI model, the natural question is when to stop improving and finally use the result. In practice, there are two common and reasonable stopping criteria. The first is when the model reaches a level of performance that is good enough for our needs. The second is when further optimization no longer brings noticeable improvement. Both ideas match the cooking analogy. Once you are satisfied with the flavor of the cake, you can simply stop adjusting the recipe and use the version you have. Alternatively, if you have the patience and the resources, you can keep experimenting as long as the attempts still produce a noticeable improvement in taste.

Finally, to truly evaluate an AI model, scientists perform a separate validation phase in which the model is tested on input features that were not used during training. This shows whether the model learned patterns that work on new examples or simply became very good at the examples it practiced on. The second case is called overfitting: the model performs well on its training data but poorly on new data. In the cooking world, this is like inviting friends to taste the cake once you are satisfied with the optimized recipe. Since they were not part of your trial and error process, their reactions reveal how well the recipe works for others.

To summarize, training an AI model is like developing a recipe through repeated experiments: each mistake provides information that can be used to improve the result. Optimization algorithms guide these improvements, while validation helps check that the model works on new examples and not only those it encountered during training. Understanding this process is important because it shows that AI does not learn by magic: its abilities depend on careful training, testing, and human decisions.

Glossary

Data: ↑ Collection of information, often in the form of numbers, text, or images.

Model: ↑ A computer program that learns patterns connecting inputs to outputs, then uses those patterns to produce outputs for new inputs.

Artificial Neural Network: ↑ A set of interconnected neurons arranged in layers that work together to model complex patterns in data.

Digitize: ↑ To convert real-world information like image into numerical form so that computers can process it.

Neuron: ↑ A computational unit that takes inputs, transforms them using its parameters, and produces an output.

Parameters: ↑ Adjustable numerical values inside a neural network that determine how neurons combine and transform inputs.

Optimization Algorithm: ↑ A step-by-step procedure used by a computer to find the best possible solution to a problem.

Gradient: ↑ Information that tells the optimization algorithm how to adjust the parameters.

Conflict of Interest

The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Acknowledgments

MS has received funding from the European Union’s Horizon 2020 research and innovation program under the Marie Sklodowska-Curie grant agreement No 101034413.

AI Tool Statement

The author(s) declared that Generative AI was used in the creation of this manuscript. Gen AI (ChatGPT and Gemini) was used in the creation of this manuscript to correct grammar and language flow of the text.

Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.

References

[1] ↑ Robbins, H., and Monro, S. 1951. A stochastic approximation method. Ann. Math. Stat. 22:400–7. doi: 10.1214/aoms/1177729586

[2] ↑ Duchi, J., Hazan, E., and Singer, Y. 2011. Adaptive subgradient methods for online learning and stochastic optimization. J. Mach. Learn. Res. 12:2121–59. doi: 10.5555/1953048.2021068

[3] ↑ Bottou, L., Curtis, F. E., and Nocedal, J. 2018. Optimization methods for large-scale machine learning. SIAM Rev. 60:223–311. doi: 10.48550/arXiv.1606.04838