The Many Applications of Gradient Descent in TensorFlow

TensorFlow is one of the leading tools for training deep learning models. Outside that space, it may seem intimidating and unnecessary, but it has many creative uses—like producing highly effective adversarial input for AI systems.

Last updated: Sep 9, 2026

Toptalauthors are vetted experts in their fields and write on topics in which they have demonstrated experience. All of our content is peer reviewed and validated by Toptal experts in the same field.

TensorFlow is one of the leading tools for training deep learning models. Outside that space, it may seem intimidating and unnecessary, but it has many creative uses—like producing highly effective adversarial input for AI systems.

Last updated: Sep 9, 2026

Toptalauthors are vetted experts in their fields and write on topics in which they have demonstrated experience. All of our content is peer reviewed and validated by Toptal experts in the same field.
Alan Reiner

Alan’s ML expertise covers visual target recognition models for missile defense systems, real-time NLP, and financial evaluation tools.

Previously At

Novetta
Share

Google’s TensorFlow is one of the leading tools for training and deploying deep learning models. It’s able to optimize wildly complex neural-network architectures with hundreds of millions of parameters, and it comes with a wide array of tools for hardware acceleration, distributed training, and production workflows. These powerful features can make it seem intimidating and unnecessary outside of the domain of deep learning.

But TensorFlow can be both accessible and usable for simpler problems not directly related to training deep learning models. At its core, TensorFlow is just an optimized library for tensor operations (vectors, matrices, etc.) and the calculus operations used to perform gradient descent on arbitrary sequences of calculations. Experienced data scientists will recognize “gradient descent” as a fundamental tool for computational mathematics, but it usually requires implementing application-specific code and equations. As we’ll see, this is where TensorFlow’s modern “automatic differentiation” architecture comes in.

TensorFlow Use Cases

Example 1: Linear Regression with Gradient Descent in TensorFlow 2.0

Example 1 Notebook

Before getting to the TensorFlow code, it’s important to be familiar with gradient descent and linear regression.

What Is Gradient Descent?

In the simplest terms, it’s a numerical technique for finding the inputs to a system of equations that minimize its output. In the context of machine learning, that system of equations is our model, the inputs are the unknown parameters of the model, and the output is a loss function to be minimized, that represents how much error there is between the model and our data. For some problems (like linear regression), there are equations to directly calculate the parameters that minimize our error, but for most practical applications, we require numerical techniques like gradient descent to arrive at a satisfactory solution.

The most important point of this article is that gradient descent usually requires laying out our equations and using calculus to derive the relationship between our loss function and our parameters. With TensorFlow (and any modern auto-differentiation tool), the calculus is handled for us, so we can focus on designing the solution, and not have to spend time on its implementation.

Here’s what it looks like on a simple linear regression problem. We have a sample of the heights (h) and weights (w) of 150 adult males, and start with an imperfect guess of the slope and standard deviation of this line. After about 15 iterations of gradient descent, we arrive at a near-optimal solution.

Two synchronized animations. The left side shows a height-weight scatterplot, with a fitted line that starts far from the data, then quickly moves toward it, slowing down before it finds the final fit. The right side shows a graph of loss versus iteration, with each frame adding a new iteration to the graph. The loss starts out above the top of the graph at 2,000, but quickly approaches the minimum loss line within a few iterations in what appears to be a logarithmic curve.

Let’s see how we produced the above solution using TensorFlow 2.0.

For linear regression, we say that weights can be predicted by a linear equation of heights.

w-subscript-i,pred equals alpha dot-product h-subscript-i plus beta.

We want to find parameters α and β (slope and intercept) that minimize the average squared error (loss) between the predictions and the true values. So our loss function (in this case, the “mean squared error,” or MSE) looks like this:

MSE equals one over N times the sum from i equals one to N of the square of the difference between w-subscript-i,true and w-subscript-i,pred.

We can see how the mean squared error looks for a couple of imperfect lines, and then with the exact solution (α=6.04, β=-230.5).

Three copies of the same height-weight scatterplot, each with a different fitted line. The first has w = 4.00 * h + -120.0 and a loss of 1057.0; the line is below the data and less steep than it. The second has w = 2.00 * h + 70.0 and a loss of 720.8; the line is near the upper part of the data points, and even less steep. The third has w = 60.4 * h + -230.5 and a loss of 127.1; the line passes through the data points such that they appear evenly clustered around it.

Let’s put this idea into action with TensorFlow. The first thing to do is code up the loss function using tensors and tf.* functions.

def calc_mean_sq_error(heights, weights, slope, intercept):
    predicted_wgts = slope * heights + intercept
    errors = predicted_wgts - weights
    mse = tf.reduce_mean(errors**2)
    return mse

This looks pretty straightforward. All the standard algebraic operators are overloaded for tensors, so we only have to make sure the variables we are optimizing are tensors, and we use tf.* methods for anything else.

Then, all we have to do is put this into a gradient descent loop:

def run_gradient_descent(
    heights,
    weights,
    init_slope,
    init_icept,
    learning_rate,
):
    # Create the variables outside the compiled loop.
    tf_slope = tf.Variable(init_slope, dtype=tf.float32)
    tf_icept = tf.Variable(init_icept, dtype=tf.float32)

    @tf.function
    def optimize():
        for _ in tf.range(25):
            with tf.GradientTape() as tape:
                predictions = tf_slope * heights + tf_icept
                errors = predictions - weights
                loss = tf.reduce_mean(errors**2)

            gradients = tape.gradient(
                loss,
                [tf_slope, tf_icept],
            )

            tf_slope.assign_sub(
                learning_rate * gradients[0]
            )
            tf_icept.assign_sub(
                learning_rate * gradients[1]
            )

    optimize()
    return tf_slope, tf_icept

The inner optimize() function uses @tf.function to compile all 25 iterations into a TensorFlow graph. tf.range() keeps the loop inside that graph, while assign_sub() updates the two variables in place. Creating the variables outside optimize() keeps variable creation out of the compiled graph and allows the graph to update their values in place.

Let’s take a moment to appreciate how neat this is. Gradient descent requires calculating derivatives of the loss function with respect to all variables we are trying to optimize. Calculus is supposed to be involved, but we didn’t actually do any of it. The magic is in the fact that:

  1. TensorFlow builds a computation graph of every calculation done under a tf.GradientTape().
  2. TensorFlow knows how to calculate the derivatives (gradients) of every operation, so that it can determine how any variable in the computation graph affects any other variable.

How does the process look from different starting points?

The same synchronized graphs as before, but also synchronized to a similar pair of graphs beneath them for comparison. The lower pair's loss-iteration graph is similar but seems to converge faster; its corresponding fitted line starts from above the data points rather than below, and closer to its final resting place.

Gradient descent gets remarkably close to the optimal MSE, but actually converges to a substantially different slope and intercept than the optimum in both examples. In some cases, this is simply gradient descent converging to local minimum, which is an inherent challenge with gradient descent algorithms. But linear regression provably only has one global minimum. So how did we end up at the wrong slope and intercept?

In this case, the issue is that we oversimplified the code for the sake of demonstration. We didn’t normalize our data, and the slope parameter has a different characteristic than the intercept parameter. Tiny changes in slope can produce massive changes in loss, while tiny changes in intercept have very little effect. This huge difference in scale of the trainable parameters leads to the slope dominating the gradient calculations, with the intercept parameter almost being ignored.

So gradient descent effectively finds the best slope very close to the initial intercept guess. And since the error is so close to optimum, the gradients around it are tiny, so each successive iteration moves only a tiny bit. Normalizing our data first would have dramatically improved this phenomenon, but it wouldn’t have eliminated it.

This was a relatively simple example, but we’ll see in the next sections that this “auto-differentiation” capability can handle some pretty complex stuff.

Applying Gradient Descent With Keras 3

Although the examples above use TensorFlow-specific APIs, Keras 3 also supports JAX and PyTorch backends. This makes it important to distinguish the gradient descent concepts that transfer across frameworks from the code used to implement them.

The gradient descent process remains consistent across all three backends: Run a forward pass, calculate the loss, derive its gradients, and update the relevant variables. Models built with Keras layers and keras.ops can generally move between them without changes to their architecture.

Low-level automatic differentiation and compilation remain backend-specific. TensorFlow uses tf.GradientTape and @tf.function; JAX provides tools such as jax.value_and_grad() and jax.jit(); and PyTorch uses backward() or torch.autograd.grad(), alongside torch.compile(). The custom loops in this article therefore remain specific to TensorFlow. Developers who need portable Keras 3 code can define the model, loss, and optimizer through Keras APIs and use built-in training methods such as fit(), which Keras adapts to the selected backend.

Example 2: Maximally Spread Unit Vectors

Example 2 Notebook

This next example is based on a fun deep learning exercise in a deep learning course I took last year.

The gist of the problem is that we have a “variational auto-encoder” (VAE) that can produce realistic faces from a set of 32 normally distributed numbers. For suspect identification, we want to use the VAE to produce a diverse set of (theoretical) faces for a witness to choose from, then narrow the search by producing more faces similar to the ones that were chosen. For this exercise, it was suggested to randomize the initial set of vectors, but I wanted to find an optimal initial state.

We can phrase the problem like this: Given a 32-dimensional space, find a set of X unit vectors that are maximally spread apart. In two dimensions, this is easy to calculate exactly. But for three dimensions (or 32 dimensions!), there is no straightforward answer. However, if we can define a proper loss function that is at its minimum when we have achieved our target state, maybe gradient descent can help us get there.

Two graphs. The left graph, Initial State for All Experiments, has a central point connected to other points, almost all of which form a semi-circle around it; one point stands roughly opposite the semi-circle. The right graph, Target State, is like a wheel, with spokes spread out evenly.

We will start with a randomized set of 20 vectors as shown above and experiment with three different loss functions, each one with increasing complexity, to demonstrate TensorFlow’s capabilities.

Let’s first define our training loop. We will put all the TensorFlow logic under the self.calc_loss() method, and then we can simply override that method for each technique, recycling this loop.

# Define the framework for trying different loss functions.
# The base class implements each optimization step, while subclasses
# override self.calc_loss().
class VectorSpreadAlgorithm:
    def __init__(self, vecs):
        # Store the vectors as a variable so the compiled function
        # can update them in place.
        self.vecs = tf.Variable(vecs, dtype=tf.float32)

    def calc_loss(self, tensor2d):
        raise NotImplementedError(
            "Define this in your derived class"
        )

    @tf.function
    def one_iter(self, learning_rate):
        with tf.GradientTape() as tape:
            loss = self.calc_loss(self.vecs)

        gradients = tape.gradient(loss, self.vecs)
        self.vecs.assign_sub(learning_rate * gradients)

        return loss

The @tf.function decorator compiles each optimization step into a TensorFlow graph. Because the surrounding visualization calls one_iter() repeatedly, TensorFlow can reuse the compiled graph instead of reconstructing the operations for every iteration. Storing self.vecs as a tf.Variable allows the function to update the vectors in place with assign_sub().

The first technique to try is the simplest. We define a spread metric that is the angle of the vectors that are closest together. We want to maximize spread, but it is conventional to make it a minimization problem. So we simply take the negative of the spread metric:

class VectorSpread_Maximize_Min_Angle(VectorSpreadAlgorithm):
    def calc_loss(self, tensor2d):
        angle_pairs = tf.acos(tensor2d @ tf.transpose(tensor2d))
        disable_diag = tf.eye(
            tf.shape(tensor2d)[0],
            dtype=tensor2d.dtype,
        ) * (2 * np.pi)
        spread_metric = tf.reduce_min(angle_pairs + disable_diag)

        # Convention is to return a quantity to be minimized, but we want
        # to maximize spread. So return negative spread
        return -spread_metric

Some Matplotlib magic will yield a visualization.

An animation going from the initial state to the target state. The lone point stays fixed, and the rest of the spokes in the semi-circle take turns jittering back and forth, slowly spreading out and not achieving equidistance even after 1,200 iterations.

This is clunky (quite literally!) but it works. Only two of the 20 vectors are updated at a time, growing the space between them until they are no longer the closest, then switching to increasing the angle between the new two-closest vectors. The important thing to notice is that it works. We see that TensorFlow was able to pass gradients through the tf.reduce_min() method and the tf.acos() method to do the right thing.

Let’s try something a bit more elaborate. We know that at the optimal solution, all vectors should have the same angle to their closest neighbors. So let’s add “variance of minimum angles” to the loss function.

class VectorSpread_MaxMinAngle_w_Variance(VectorSpreadAlgorithm):
    def spread_metric(self, tensor2d):
        """Assumes all rows are already normalized."""
        angle_pairs = tf.acos(tensor2d @ tf.transpose(tensor2d))
        disable_diag = tf.eye(
            tf.shape(tensor2d)[0],
            dtype=tensor2d.dtype,
        ) * (2 * np.pi)
        all_mins = tf.reduce_min(
            angle_pairs + disable_diag,
            axis=1,
        )

        # Same calculation as before: find the min-min angle
        min_min = tf.reduce_min(all_mins)

        # But now also calculate the variance of the min angles vector
        avg_min = tf.reduce_mean(all_mins)
        var_min = tf.reduce_sum(tf.square(all_mins - avg_min))

        # Our spread metric now includes a term to minimize variance
        spread_metric = min_min - 0.4 * var_min

        # As before, want negative spread to keep it a minimization problem
        return -spread_metric

An animation going from the initial state to the target state. The lone spoke does not stay fixed, quickly moving around toward the rest of the spokes in the semi-circle; instead of closing two gaps either side the lone spoke, the jittering now closes one large gap over time. Equidistance is here also not quite achieved after 1,200 iterations.

That lone northward vector now rapidly joins its peers, because the angle to its closest neighbor is huge and spikes the variance term which is now being minimized. But it’s still ultimately driven by the globally-minimum angle which remains slow to ramp up. Ideas I have to improve this generally work in this 2D case, but not in any higher dimensions.

But focusing too much on the quality of this mathematical attempt is missing the point. Look at how many tensor operations are involved in the mean and variance calculations, and how TensorFlow successfully tracks and differentiates every computation for every component in the input matrix. And we didn’t have to do any manual calculus. We just threw some simple math together, and TensorFlow did the calculus for us.

Finally, let’s try one more thing: a force-based solution. Imagine that every vector is a small planet tethered to a central point. Each planet emits a force that repels it from the other planets. If we were to run a physics simulation of this model, we should end up at our desired solution.

My hypothesis is that gradient descent should work, too. At the optimal solution, the tangent force on every planet from every other planet should cancel out to a net zero force (if it weren’t zero, the planets would be moving). So let’s calculate the magnitude of force on every vector and use gradient descent to push it toward zero.

First, we need to define the method that calculates force using tf.* methods:

class VectorSpread_Force(VectorSpreadAlgorithm):
    
    def force_a_onto_b(self, vec_a, vec_b):
        # Calc force assuming vec_b is constrained to the unit sphere
        diff = vec_b - vec_a
        norm = tf.sqrt(tf.reduce_sum(diff**2))
        unit_force_dir = diff / norm
        force_magnitude = 1 / norm**2
        force_vec = unit_force_dir * force_magnitude

        # Project force onto this vec, calculate how much is radial
        b_dot_f = tf.tensordot(vec_b, force_vec, axes=1)
        b_dot_b = tf.tensordot(vec_b, vec_b, axes=1)
        radial_component =  (b_dot_f / b_dot_b) * vec_b

        # Subtract radial component and return result
        return force_vec - radial_component

Then, we define our loss function using the force function above. We accumulate the net force on each vector and calculate its magnitude. At our optimal solution, all forces should cancel out and we should have zero force.

def calc_loss(self, tensor2d):
    n_vec = tensor2d.shape[0]
    all_force_list = []

    for this_idx in range(n_vec):

        # Accumulate force of all other vecs onto this one
        this_force_list = []
        for other_idx in range(n_vec):

            if this_idx == other_idx:
                continue

            this_vec = tensor2d[this_idx, :]
            other_vec = tensor2d[other_idx, :]

            tangent_force_vec = self.force_a_onto_b(other_vec, this_vec)
            this_force_list.append(tangent_force_vec)

        # Use list of all N-dimensional force vecs. Stack and sum.
        sum_tangent_forces = tf.reduce_sum(tf.stack(this_force_list))
        this_force_mag = tf.sqrt(tf.reduce_sum(sum_tangent_forces**2))

        # Accumulate all magnitudes, should all be zero at optimal solution
        all_force_list.append(this_force_mag)

    # We want to minimize total force sum, so simply stack, sum, return
    return tf.reduce_sum(tf.stack(all_force_list))

An animation going from the initial state to the target state. The first few frames see rapid movement in all spokes, and after only 200 iterations or so, the overall picture is already fairly close to the target. Only 700 iterations are shown in total; after the 300th, angles are changing only minutely with each frame.

Not only does the solution work beautifully (besides some chaos in the first few frames), but the real credit goes to TensorFlow. This solution involved multiple for loops, an if statement, and a huge web of calculations, and TensorFlow successfully traced gradients through all of it for us.

Example 3: Generating Adversarial AI Inputs

At this point, readers may be thinking, "Hey! This post wasn't supposed to be about deep learning!" However, we’re not training a model here. We’re using gradient descent to alter the input to a pretrained image classifier while leaving its parameters unchanged. For this example, we’ll attack a Vision Transformer (ViT), an architecture that divides an image into patches and processes them through Transformer encoder layers. We’ll use the ViT-B/16 model available through KerasHub, pretrained on the 1,000-class ImageNet dataset. The model accepts 224 x 224 images, divides them into 16 x 16 patches, and contains approximately 85.8 million parameters.

import keras
import keras_hub
import tensorflow as tf

model = keras_hub.models.ViTImageClassifier.from_preset(
    "vit_base_patch16_224_imagenet"
)
model.trainable = False

KerasHub packages the required preprocessing with the classifier, allowing the model to accept image tensors containing pixel values between 0 and 255. Its output contains a probability for each ImageNet class.

We can load an image and convert it into the required batch tensor:

image = keras.utils.load_img(
    "sample_image.jpg",
    target_size=(224, 224),
)
original_img = keras.utils.img_to_array(image)
original_img = tf.expand_dims(original_img, axis=0)

The goal is to change the model’s prediction without making the image appear meaningfully different to a human observer. To do that, we need to determine how a small change to each pixel would affect the classification loss.

During training, gradient descent calculates how the model’s parameters affect its loss and updates those parameters. Here, the model remains fixed. We calculate the gradient of the loss with respect to the input image and update its pixels instead.

The following function performs an iterative fast gradient sign method attack. Each iteration moves the pixels in the direction that increases the loss for the original class. It also limits the total perturbation to four intensity levels per pixel, preserving the visual similarity between the original and modified images.

@tf.function
def adversarial_modify(
    original_img,
    original_label,
    num_steps=4,
    step_size=1.0,
    epsilon=4.0,
):
    adversarial_img = tf.identity(original_img)

    for _ in tf.range(num_steps):
        with tf.GradientTape() as tape:
            tape.watch(adversarial_img)

            predictions = model(adversarial_img, training=False)
            loss = keras.losses.sparse_categorical_crossentropy(
                original_label,
                predictions,
            )

        gradient = tape.gradient(loss, adversarial_img)

        # Increase the loss assigned to the original class.
        adversarial_img += step_size * tf.sign(gradient)

        # Limit the total change to the permitted pixel range.
        perturbation = tf.clip_by_value(
            adversarial_img - original_img,
            -epsilon,
            epsilon,
        )
        adversarial_img = tf.clip_by_value(
            original_img + perturbation,
            0.0,
            255.0,
        )

    return adversarial_img

This is an untargeted white-box attack: It assumes access to the model and attempts to make the original prediction incorrect without prescribing which class should replace it.

We can identify the model’s original prediction and pass that class to the attack:

original_predictions = model(original_img, training=False)
original_label = tf.argmax(
    original_predictions,
    axis=-1,
    output_type=tf.int32,
)

adversarial_img = adversarial_modify(
    original_img,
    original_label,
)

The attack changes each pixel by no more than four intensity levels. These small adjustments may be difficult for a person to detect, yet they can alter the model’s classification because the gradient coordinates thousands of pixel changes toward the same objective.

Switching from a convolutional neural network to a ViT does not change the underlying gradient descent principle. A ViT uses attention to process image patches, but its output remains differentiable with respect to its input. tf.GradientTape can therefore trace the calculation back to the image and determine how each pixel should be adjusted.

Final Thoughts: Gradient Descent Optimization

Throughout this article, we’ve looked at manually applying gradients to our trainable parameters for the sake of simplicity and transparency. However, in the real world, data scientists should jump right into using optimizers, because they tend to be much more effective, without adding any code bloat.

There are many popular optimizers, including RMSprop, Adagrad, and Adadelta, but the most common is probably Adam. Sometimes, they are called “adaptive learning rate methods” because they dynamically maintain a different learning rate for each parameter. Many of them use momentum terms and approximate higher-order derivatives, with the goal of escaping local minima and achieving faster convergence.

In an animation borrowed from Sebastian Ruder, we can see the path of various optimizers descending a loss surface. The manual techniques we have demonstrated are most comparable to “SGD.” The best-performing optimizer won’t be the same one for every loss surface; however, more advanced optimizers do typically perform better than the simpler ones.

An animated contour map, showing the path taken by six different methods to converge on a target point. SGD is by far the slowest, taking a steady curve from its starting point. Momentum initially goes away from the target, then criss-crosses its own path twice before heading toward it not entirely directly, and seeming to overshoot it and then backtrack. NAG is similar, but doesn't stray quite as far from the target and criss-crosses itself only once, generally reaching the target faster and overshooting it less. Adagrad starts off in a straight line that's the most off-course, but very quickly does a hair-pin turn toward the hill the target is on, and curving toward it faster than the first three. Adadelta has a similar path, but with a smoother curve; it overtakes Adagrad and stays ahead of it after the first second or so. Finally, Rmsprop follows a very similar path to Adadelta, but leans slightly closer to the target early on; notably, its course is much more steady, making it lag behind Adagrad and Adadelta for most of the animation; unlike the other five, it seems to have two sudden, rapid jumps in two different directions near the end of the animation before ceasing movement, while the others, in the last moment, continue to slowly creep along by the target.

However, it is rarely useful to be an expert on optimizers—even for those keen on providing artificial intelligence development services. It is a better use of developers’ time to familiarize themselves with a couple, just to understand how they improve gradient descent in TensorFlow. After that, they can just use Adam by default and try different ones only if their models aren’t converging.

For readers who are really interested in how and why these optimizers work, Ruder’s overview—in which the animation appears—is one of the best and most exhaustive resources on the topic.

Let’s update our linear regression solution from the first section to use optimizers. The following is the original gradient descent code using manual gradients.

# Manual gradient descent operations
def run_gradient_descent(
    heights,
    weights,
    init_slope,
    init_icept,
    learning_rate,
):
    tf_slope = tf.Variable(
        init_slope,
        dtype=tf.float32,
    )
    tf_icept = tf.Variable(
        init_icept,
        dtype=tf.float32,
    )

    @tf.function
    def optimize():
        for _ in tf.range(25):
            with tf.GradientTape() as tape:
                predictions = (
                    tf_slope * heights + tf_icept
                )
                errors = predictions - weights
                loss = tf.reduce_mean(errors**2)

            gradients = tape.gradient(
                loss,
                [tf_slope, tf_icept],
            )

            tf_slope.assign_sub(
                learning_rate * gradients[0]
            )
            tf_icept.assign_sub(
                learning_rate * gradients[1]
            )

    optimize()
    return tf_slope, tf_icept

Now, here is the same code using an optimizer instead. You will see that it’s hardly any extra code (changed lines are highlighted in blue):

# Gradient descent with an RMSprop optimizer
def run_gradient_descent(
    heights,
    weights,
    init_slope,
    init_icept,
    learning_rate,
):
    tf_slope = tf.Variable(
        init_slope,
        dtype=tf.float32,
    )
    tf_icept = tf.Variable(
        init_icept,
        dtype=tf.float32,
    )

    # Group the trainable parameters into a list.
    trainable_params = [tf_slope, tf_icept]

    # Define the optimizer outside the compiled loop.
    optimizer = keras.optimizers.RMSprop(
        learning_rate=learning_rate
    )

    @tf.function
    def optimize():
        for _ in tf.range(25):
            with tf.GradientTape() as tape:
                predictions = (
                    tf_slope * heights + tf_icept
                )
                errors = predictions - weights
                loss = tf.reduce_mean(errors**2)

            gradients = tape.gradient(
                loss,
                trainable_params,
            )

            optimizer.apply_gradients(
                zip(gradients, trainable_params)
            )

    optimize()
    return tf_slope, tf_icept

That’s it! We defined an RMSprop optimizer outside of the gradient descent loop, and then we used the optimizer.apply_gradients() method after each gradient calculation to update the trainable parameters. The optimizer is defined outside of the loop because it will keep track of historical gradients for calculating extra terms like momentum and higher-order derivatives.

Let’s see how it looks with the RMSprop optimizer.

Similar to the previous synchronized pairs of animations; the fitted line starts above its resting place. The loss graph shows it nearly converging after a mere five iterations.

Looks great! Now let’s try it with the Adam optimizer.

Another synchronized scatterplot and corresponding loss graph animation. The loss graph stands out from the others in that it doesn't strictly continue to get closer to the minimum; instead, it resembles the path of a bouncing ball. The corresponding fitted line on the scatterplot starts above the sample points, swings toward the bottom of them, then back up but not as high, and so on, with each change of direction being closer to a central position.

Whoa, what happened here? It appears the momentum mechanics in Adam cause it to overshoot the optimal solution and reverse course multiple times. Normally, this momentum mechanic helps with complex loss surfaces, but it hurts us in this simple case. This emphasizes the advice to make the choice of optimizer one of the hyperparameters to tune when training your model.

Anyone wanting to explore deep learning will want to become familiar with this pattern, as it is used extensively in custom TensorFlow architectures, where there’s a need to have complex loss mechanics that are not easily wrapped up in the standard workflow. In this simple TensorFlow gradient descent example, there were only two trainable parameters, but it is necessary when working with architectures containing hundreds of millions of parameters to optimize.

Gradient Descent in TensorFlow: From Finding Minimums to Attacking AI Systems

The corresponding GitHub repository contains the original notebooks, visualizations, and inline documentation for readers who want to explore the examples in greater detail. The code in this article has been updated to reflect current TensorFlow and Keras practices, so some snippets differ from the repository versions.

I hope this article was insightful and it got you thinking about ways to use gradient descent in TensorFlow. Even if you don’t use it yourself, it hopefully makes it clearer how all modern neural network architectures work—create a model, define a loss function, and use gradient descent to fit the model to your dataset.


Google Cloud Partner badge.

As a Google Cloud Partner, Toptal’s Google-certified experts are available to companies on demand for their most important projects.

Understanding the basics

  • TensorFlow is typically used for training and deploying AI agents for a variety of applications, such as computer vision and natural language processing (NLP). Under the hood, it’s a powerful library for optimizing massive computational graphs, which is how deep neural networks are defined and trained.

  • Tensorflow is a deep learning framework created by Google for both cutting-edge AI research as well as deployment of AI applications at scale. Under the hood, it is an optimized library for doing tensor calculations and tracking gradients through them for the purposes of applying gradient descent algorithms.

  • Gradient descent is a calculus-based numerical technique used to optimize machine learning models. The error of a given model is defined as a function of the model’s parameters, and gradient descent is applied to adjust those parameters to minimize that error.

  • Gradient descent works by representing a model’s error as a function of its parameters. Using calculus, one can compute how this error would change in response to adjustments in each parameter—its gradient—then adjust those parameters iteratively until the error of the model is minimized.

  • Gradient descent is a numerical technique to find approximate minimal values of a function. It is most commonly associated with training artificial neural networks where the goal is to minimize an error or loss function.

Hire a Toptal expert on this topic.
Hire Now
Alan Reiner

Alan Reiner

Columbia, MD, United States

Member since February 14, 2020

About the author

Alan’s ML expertise covers visual target recognition models for missile defense systems, real-time NLP, and financial evaluation tools.

authors are vetted experts in their fields and write on topics in which they have demonstrated experience. All of our content is peer reviewed and validated by Toptal experts in the same field.
PREVIOUSLY AT
Novetta

World-class articles, delivered weekly.

By entering your email, you are agreeing to our privacy policy.

World-class articles, delivered weekly.

By entering your email, you are agreeing to our privacy policy.

Join the Toptal® community.