Decision Intelligence for Enterprise SEO: Presented at Tech SEO Connect

Below are the recorded live-stream and individual slides for my talk at Tech SEO Connect. Decision Intelligence is a new, exciting field that focuses on merging analytics, machine learning, data science, and other disciplines to help make better decisions with data.

“Decision Intelligence for Enterprise SEO” Slides

Embedded below is the presentation I shared at Tech SEO connect. If you have any questions or would like to discuss any of my points within the deck, feel free to reach out via LinkedIn!

Summary of Talk

This presentation introduces the DECIDE framework for applying Decision Intelligence to enterprise SEO. Decision Intelligence merges analytics, machine learning, and data science to transform raw data into informed, actionable decisions rather than relying on instinct or incomplete analysis.

The DECIDE framework covers six key steps:

D – Define your primary goal: Establish clear, actionable objectives free from bias (confirmation, selection, outcome, or recency bias). Use formal analysis types (descriptive, diagnostic, prescriptive, predictive) to frame your questions.

E – Extract data: Pull from reliable sources (BigQuery, CMS databases, log files, GSC/GA4 APIs) using proper ETL workflows. Avoid low-quality data sources and consider stratified sampling for large sites.

C – Clean & Transform: Handle duplicates, enrich data, and use visualization techniques like distributions, Lorenz curves, percentiles, and violin plots to understand data patterns and identify opportunities.

I – Integrate & Classify: Merge multiple data sources (GSC + GA4 + revenue data) to reveal true priorities. Apply performance labels, correlations, and weighted scoring to distinguish high-impact opportunities from noise.

D – Distribute insights: Convert processed data into clear, stakeholder-friendly insights using conditional summaries, intent labeling, and storytelling that connects to business goals.

E – Execute & measure: Establish benchmarks, implement changes, annotate events, and measure outcomes for at least three months while tracking external factors.

The framework emphasizes moving beyond checklist audits to data-driven prioritization, helping SEOs identify what will truly move the needle and confirm whether decisions were correct.

Presentation Transcript

Decision Intelligence for Enterprise SEO

Tech SEO Connect Conference – Full Transcript

Introduction

What’s up? All right, let’s get into it. And forewarning, I have a lot of short slides, so I’m going to be clicking through those kind of fast, so it might feel like a little bit of an EDM show while we’re here, but we’ll get through it.

So, Mike already kind of gave me an introduction. I’m a tech director, web developer, data science enthusiast for like the past four years. You can see by my GitHub credentials I be pushing code. And I’m also a musician. You can see maybe that’s a picture of me jumping like 3 feet in the air in LA. If you want to know more on how to do that, you can talk to me afterward.

Meet Toby: The Problem

But let’s meet Toby here. He’s an SEO consultant, uses all the normal tools, but he’s been a little bit disappointed with his results from some of his recommendations. Or maybe some people are not happy with what’s happening outside of his team. He feels stuck. What will move the needle? There’s so many things we can do. What’s going to be the most impactful, right?

Once he decides, how does he know he made the right decision? Like what can he do to confirm that this decision is the best out of the other options that he has?

What is Decision Intelligence?

And that is going to be decision intelligence.

Decision intelligence is a discipline—I will say across various industries it can mean and represent different things. So when I’m talking about decision intelligence for SEO, that can be a subsegment of the larger concept and discipline of decision intelligence. So if you go and research this topic, there’s a lot of breadth and depth to it. But it’s basically the difference between making poor decisions and making really informed decisions—not just data-driven decisions, but really honing in on data and using it the best way that you can.

And it’s taking your raw data and helping you ask and answer specific questions: What happened? Why is it happening? How to improve outcomes? To lead you to the best options and the best decisions that you can make with all available data that you have.

Is Your Organization Ready?

And before I actually get into the framework and the meat of the discussion, these are some questions and points that you should be asking. Is your organization even ready for something like decision intelligence? And spoiler, these are great questions to ask potential clients that you’re going to onboard because this can really vet how sophisticated their processes are.

Key Questions:

  • What is their willingness to change? Can they pivot? Are they willing to pivot when something contradicts their assumptions?
  • What is their current process? How are their decisions made? I’m going to say most case: instinct, a larger committee of leadership, or data. I’m going to probably say it’s probably largely the first two unfortunately.
  • What are their development priorities? Can their development team respond in a timely manner? What do they find valuable with their time?
  • How are they measuring their decisions? Are they actually doing that accurately?

Great questions to ask and kind of a precursor to if an organization is ready for this.

The Outcome vs. Process Debate

And there’s a survey from Gartner on decision intelligence where half the people said that good decisions were defined by the outcome—the end—while the other 50% said it was defined by its process. And I think that’s really interesting because not all decisions are always going to lead to good outcomes. But if you have a really, really good defined process, you’re in the best position where you’re always going to—maybe in most cases—have the best possible option even if it doesn’t lead to the best expected outcomes.


The DECIDE Framework

So we have the DECIDE model here which is going to bring us through some steps and a framework that I personally use subconsciously and now it’s real. But we’re going to go through each step and kind of what’s involved in this process to bring you to the best results you can get from using data.


D – Define Your Primary Goal

This is the most important step. And this is where probably a lot of SEOs can get stuck or just maybe be a little bit lazy with it. What does success look like for an analysis?

When you’re trying to figure something out, it’s the computer science function: f(x) = y. You have your inputs with your data. It goes through a process and you have your outputs which should be your goal.

Key Questions for Your Goal:

  • Is there significance? Do you have enough data? Do you have access to the right data?
  • Is it connected to a business goal and not vanity?
  • Is it clear? Can you tell a story about the outcome? You’re not just sharing data and saying, “Hey, figure it out.”
  • Is it actionable? Will this analysis enable a decision?

Types of Analysis

There are common analyses that you can use to help define your goals. These are the four most commonly used in marketing:

  1. Descriptive – What happened previously?
  2. Diagnostic – Why did something happen?
  3. Prescriptive – What should we do with this data?
  4. Predictive – What will happen if we do something?

For the most part, you’re probably going to be bucketed into the first two—descriptive and diagnostic. And then there’s more advanced where you’re doing machine learning territory.

Avoiding Bias

And when you’re defining your goals, this is one of the biggest things: bias. So avoid that for real, for real on God. Bias in, bias out.

I understand, you know, even some of the best, most well-intentioned SEOs, marketing analysts, whatever, are susceptible to this. And honestly, I think it’s pretty harmful because we’ve all seen those reports that are really cherry-picked. Like, you know, when someone screenshots like one fraction of a GSC chart and they’re like, “Look at this growth.” But it definitely harms our industry reputation because if you see that and you’re just getting burned, those are the kinds of things that actually can impact someone’s perception of the industry.

Common Biases to Avoid:

Confirmation Bias – Seeking data that supports your hypothesis while ignoring evidence. You have the freaking blinders on. This is also very common with a lot of junior SEOs, cherry-picking things, being dishonest about what you’re including and excluding. Just be honest. Use all available data and be a trustworthy partner to whoever you’re working with and be honest with what you’re looking at.

Selection Bias – Analyzing only the successful or visible data while missing the complete picture. So shocker: if you’re using GSC and you’re just exporting a thousand rows, you’re missing so much data it’s not even funny. Using SEO tools for first-party research usually not the best option, or sample GA4 data.

Outcome Bias – Defining success criteria after the results. So this is basically you don’t like your analysis, things are not looking great and you’re changing it or you’re trying to remove points to make it more favorable. This is very common as well.

Recency Bias – Overweighing recent data while ignoring patterns. Clients love this. We love this because once there’s a spike, we’re like, “Oh my god, it’s changed forever.” But you kind of need to back up and say, “Okay, what was the baseline?” We always hit that emotional peak. We get those messages on Slack at like, you know, the second we clock in and then things go back to normal and it’s like, okay, it’s interesting to analyze those things, but don’t make an entire shift just because one of those events occurs.

Start with a Solid Hypothesis

And just start with a solid hypothesis: If X then Y. We want to measure success by this. We want to analyze the relationship between this and that. You can come up with some pretty interesting things. And there’s obviously a lot of domain expertise involved in this, but it can be just as simple as this.

Examples of Defining Complex Goals

Example 1: Crawled Not Indexed

Crawled not indexed—everybody knows that’s like one of the most black box things there is. And what we did with this was we defined the goal: okay, we know these pages should be indexed. They’re indexable. We’re submitting them. We want to analyze correlations of the content signals and indexation.

And what that led to was we discovered specific attributes that mattered. So in this case it was using unique images versus placeholders, the stock status and the availability. Those all had common patterns of those pages being indexed. And yeah, the others that we looked at like reviews and specifications and similarity—they weren’t as strong.

Example 2: Traffic Erosion

Another one was: why is our traffic eroding and where? I mean it’s very common. So we want to look at keyword intent groupings and analyze trends. And that output was, you can see, there’s about 48,000 keywords and we grouped those into like 20 to 30 intents. And then we ran a decay analysis on that. You can see there’s some obvious intents that are a little more volatile than others.


E – Extract Your Data

So this is part of the ETL process, which is a data science process, just to kind of formalize the process and lead you to better analyses. You got to respect the ETL workflow: extracting your data, pulling it from various sources, transforming, cleaning, joining, enriching, scoring, and then loading, which is basically your end result of like either a dashboard or visualization or something you can kind of maybe tell a story with.

Data Source Quality Tiers

Godly Tier:

  • BigQuery
  • Your CMS database
  • Your log files
  • APIs like GSC and GA4

Very Solid:

  • SERPs APIs
  • Web scraper

Fine (for research, not first-party analysis):

  • Ahrefs
  • SEMrush
  • Other platforms

They’re great for research, not as reliable for analyzing your first-party data or your traffic performance.

Avoid for Data Analysis:

  • The UIs of GSC or GA – whatever, they’re fine, I would never use these sources for getting data obviously. And if you are, you got to stop right now.

Sampling for Large Sites

For large sites you don’t need to analyze everything. You could use random sampling. And this is a whole other discussion, but you can use stratified sampling. It’s basically taking your data set—could be your website, your pages, queries—breaking it down to strata, random selection, and then sampling. And it gives you a distribution of your site without having to analyze like 800,000 pages or 200 million pages depending on how big your site is.


C – Clean & Transform Your Data

Moving on to transform. Well, this is actually the second part of this process. Cleaning your data: handling any duplicates, nulls, filtering, enriching, categorizing, labeling, and then joining—just merging various sources by URL, keyword, or date.

And transforming data can help you better understand your data obviously.

Visualization Techniques

Distributions – A good way to visualize this aspect. So you can identify outliers, anomalies, and just understand the skew and the weights of your data.

Example: We’re looking at the distribution of days since last crawled by indexation. And you can see that there’s a greater concentration of pages that were indexed when they’re more recently crawled versus pages that were—you can see the 13-day rule coming into play here. Just a good way to see how the weights are. So this should tell us, okay, we need to get more pages crawled more frequently because more pages are indexed that are crawled more recently.

Lorenz Curve – A cumulative distribution, which I use this probably most times when I’m looking at getting a representative sample of data, which is a prioritized sample of your data.

Example: What you’re seeing here is you’re sampling the 70th percentile of pages because they account for 85% of your traffic. So why do you need to even look at the remaining portion if you’re really trying to hone in on the pages that perform the best for your site?

Percentiles – Ways to identify natural break points within your data. This is really interesting for site audits. If you’re looking at how a site’s built up, quartile bins, grouping data based on the performance segment to help you understand the weights of your data.

Example: You can see you can almost always categorize pages at zero clicks into their own bucket. It’s likely noise for your analysis. And then you have your buckets of the 25th percentile to the 90th depending on what you’re prioritizing. Whether you want to prioritize the 25th percentile or the 90th percentile, which would be the high performers, you at least know how those pages are divided up.

Violin Plots/Box Plots – You can use these to actually visualize the spread and variance across different metrics to identify opportunities.

Example: We looked at internal linking distribution by site segment. And you can see that the median internal links for segment one—since it’s not as long—were lower than the other segments. That was actually the most important segment of the site. So there’s a clear opportunity that we need to have more internal links for those pages in segment one.

Data Labeling

Data labeling – assigning meaningful categories to unstructured data could transfer raw data into analyzable segments. And this really just shows more meaningful insights when you group things together. You know, you kind of bring it all together and condense it. You’re not looking at individual rows of data anymore.

Performance Labels – A great way to do this. And you can do this many different ways, but it’s dividing pages into segments based on how they perform.

Example: We had four click bins here, which is the number of clicks pages got. And what was interesting is that over time in this example, there were more pages that had higher click performance that were going down in indexation and then there were pages that had lower click performance that were going up in indexation. So it’s a way to see an interesting problem happen when you look at it from this perspective.

Rule-Based Category Groups – This is a little more complicated. Using metric groups and category groups to understand different performance across your different categories.

Example: This was basically looking at when we actually have categories on a site, we can create crawl buckets based off the recency of the crawl. And you can see from the left to the right, the gray means that that category was not crawled. So like the first two categories, almost 80% of those pages in that category haven’t been crawled. So that’s a clear opportunity when going down that scale in descending order that there’s some categories that definitely need to be having some attention.

Correlations

Correlations – use these cautiously. It’s helpful to understand relationships between variables, but some things move together more obviously.

Multi-dimensional Correlations – Looking at like discovery crawls versus refresh crawls versus indexation. There was a positive correlation. I mean this p-value is very debated in the data science community. So I’m just going to keep it very simple and keep that, you know, if it’s below 0.05 it suggests that there’s significance. But what we found here was that discovery crawls and indexation specifically were really closely related in correlation.

Be Skeptical – And be skeptical of correlations. They occur but they’re good for hypothesis testing, but just because there’s a strong correlation doesn’t equal causation. Obviously.

Examples:

  • Make sense: Index, clicks, impressions. We know that impressions and indexation move very closely together. Sometimes clicks, it’s not as strong, but you can see the correlations here.
  • Don’t make sense but are interesting: I’m really not going to go into detail on this one. It was looking at sitemap index order and URL order within the sitemaps and then the crawl frequency. It was inconclusive, but you know, you just never know sometimes.

I – Integrate & Classify Priority Levels

Connecting your business value with data and working towards what matters. Priority can be risks like impact of loss, decay of a metric, instability, threshold breach, looking at crawl budget inefficiency. And then there’s opportunities things—things like potential for gain, underperforming assets, growth trajectory, optimizing for upward momentum, efficiency gaps, underutilized assets. So like things like good positions but low clicks.

Priority Examples

Example 1: Crawl Budget Optimization

This is an example of using where we were looking at potential for gains within crawl budget. So we saw that Bing was crawling very aggressively in relation to Googlebot. And in this client specifically, the value of Googlebot’s indexation within Google was much more valuable than Bing. So we applied a crawl delay. And then you can see the clear trade-off here where the crawling of Googlebot greatly improved and that led to a higher brand exposure which was their business value in this case.

Example 2: Navbar Risk Analysis

Also running a risk analysis for a navbar. So in this case would be optimizing a navbar. This client wanted us to potentially remove or change things within their navbar. And we basically looked at engagement, link equity, and revenue attribution to find items within those categories within their nav that were high risk and low risk.

So you can see there’s specific categories that are really strongly purple, meaning like, okay, we shouldn’t touch those. They have a lot of those combined metrics. And then other categories like six and nine where there were some lower risk items where we wouldn’t really have as much impact on the site.

Example 3: Content Decay Analysis

Content decay – this is also very common, you know, because period over period, month over month, year over year, you’re missing the entire context of data in between those points. So if you’re analyzing just two points in time, you’re missing an entire piece of context.

So metric slopes are much more meaningful and actually give you true representation of pages that are either sloping upward or downward. Year-over-year or period over period are fine, but there’s a lot missing there.

Integrating Multiple Data Sources

You might be solving the wrong problem perfectly. We do this all the time where, you know, Core Vitals for example…

Merging various data sources allows priority to be even more clear.

Single Source Problems:

  • Search Console alone: High impressions—okay what does that mean? Is it a priority? What’s it doing?
  • GA4 alone: This page has engagement, is it valuable? What do we do with it?
  • Revenue data alone: It sells well, but how is SEO involved in that process?

Not getting the whole picture there.

Integrated View – You know, just combining GSC, GA4, and conversion data—highly valuable. There’s so many different ways you can cross-reference different changes in the metrics between the data sources to come up with different opportunities. Crawl logs and GA4 AI referrals looking at maybe optimize internal linking based off of what bots are crawling and what sort of referrals you’re getting from AI sources. That’s the only time I mentioned AI.

Efficiency Gaps – So this is where we’re looking at search console, GA4, and conversions. And you can see that roughly maybe 70% of their revenue was coming from a source that wasn’t SEO. So we can ask the question of like okay, maybe SEO can play a larger role in revenue. Maybe there’s something we need to prioritize differently within our products.

Weighted Averages & Scoring

You can use different weighted averages. This is good for impact and priority scoring. This is just one example. You can use your own weights and your own variables there.

I’m going to skip through this one. This is actually something that I have that’s open source on my GitHub. It’s a way to take a Screaming Frog crawl with GA4, GSC and actually get an impact score where higher impact, lower impact—just allows you to have more direction on your audits and stop having those checklists.

Priority Scoring – Very simple example but you can create your own proprietary weights with this. But this is a good way to actually get sort of like, okay, we know the impact—what’s the priority with the business, with the operation?

So you can label things low, medium, high, provide a number to those values, urgency, capacity with development, and have a priority score based on that. So you can create your own weights to do this to be much more sophisticated. This is just simple for example’s sake.

The Priority Threshold

And up until this point we are up to this part where you have basically from the start: someone asks you a question, there’s an external event, something’s happening, you’re looking at data, you go through all the steps and if you aren’t passing your priority score—if it’s not greater than a certain threshold—maybe that’s something that needs to go on the backlog. If it does, then maybe we need to move forward further with the process.


D – Distribute Your Data

So the fifth, second to last step here would be distributing your data, which is part of the load of the ETL process. And stakeholders are more likely to misinterpret your data. I mean that’s just kind of how it is sometimes. They give up, have biases.

So convert your data into actual insights. You have your raw data, goes through the processing, it gets organized post-processing. Then you have your insightful data.

Programmatic Insights

You can do this programmatically. I mean, there’s a little bit of developer knowledge here, but there are some interesting ways you can do this based off of patterns within your data.

Conditional-Based Summaries – This was an example of where we took canonical patterns that were happening within the actual strings of the URLs and created unique labeling for each of those instances. And it’s much better than showing just simple examples or maybe it’s complex data patterns and you just want to get the message across of what’s happening. And in this case it was local canonicaling to non-local pages which were 76% of this issue here.

Conditional-Based Summaries and Filtering – This was where we’re looking at different slopes of pages and we sectioned out the pages that were having negative performance from a click and position perspective. Siloed those out and those are the pages—that’s your target data set for prioritizing your content and possibly refreshing.

Conditional-Based Summaries and Intent Labeling – So basically isolating your keywords per decaying page, labeling the keyword intents, showing the trends, and then kind of overlaying with anything like a Google update. And now we can see that there’s a high value intent there that we know is declining. So we already know what the pages are, the keywords are, we can then optimize and maybe hopefully address what’s happening there.

Making Data Accessible

Since we need to be using spreadsheets, always have a summary column. I mean, I deleted Excel like two years ago and told myself I’d never use it again, but I still have to use it unfortunately. So put a summary column wherever you can and help those unfortunate souls.

And speak towards the interests of your stakeholders best you can. Data as the foundation, a little bit of analytics know-how—that is going to go a long way. Storytelling, being able to tell a story with your data. And then the actual expertise of your client or your business or your site.


E – Execute & Measure

And then the last and final step here since we’re coming to the end is: you made it. If you don’t have next steps, what’s the point? You know, you have your analysis. Okay, cool. But what are you going to do with it? It needs to lead to something.

Possible Next Steps:

  • Sharing a new opportunity
  • Updating content
  • Removing pages
  • Sharing a hypothesis
  • Clarifying concern

These are all important things to share. Just because something is maybe not going to lead to the best decision you can share, there’s still information worth sharing with a really well planned out analysis.

The Execution Process

  1. Establish your benchmark
  2. Implement
  3. Annotate
  4. Measure for at least three months
  5. Report on the outcomes

Note any sort of external changes or factors like Google updates, I guess no AI updates, everything that’s happening in the world, whatever you’re into.

Keep a running log of everything that happens on a site. I do this and then I kind of remove the ones that I don’t want and some that I do. But you can have site updates with Google updates and kind of keep track of what’s happening with your site.


Conclusion

So start feeling more better about your decisions with data and use my framework. Don’t use it.

Thanks.


Posted

in

by

Comments

Leave a Reply

Discover more from Tyler Gargula

Subscribe now to keep reading and get access to the full archive.

Continue reading