Data 101: The Preeminence of the Semantic Layer

Transcript
Alrighty. Let's get going. So welcome. You are here, at an Interworks webinar. Today, we're gonna be talking about some foundational concepts when it comes to the modern data, environment data, efforts data, platforms, whatever you wanna call it. And that's the importance, the preeminence, the rise of the semantic layer. That was not always the case. So we'll do a little bit of a retrospective looking back, and then we'll take a look at a current inventory of how things are looking. And we might even look a little bit in the, looking forward. So just a little bit, I'm Robert Curtis. I'm the managing director of Interworks based out of Melbourne. I look after Asia Pacific. I've been with Interworks for about twenty years in total. So I've been here for a long, long time, and I've gotten to work with a lot of customers all over the world all different sizes and all different verticals. It's been a pleasure to get to help them. But the good news, the good result of that is is that I get a chance to, have a lot of different perspectives, a lot of different stories in terms of how other people have solved problems. So I'm happy to give you as much as I can in whatever I've learned. This series of webinars that we're doing over the next couple of weeks is is part of an extension of our effort, our booth that we had at the Snowflake World Tour, which was in Sydney a couple weeks ago. We have a white paper that is out. You can find that on interworks dot com specifically about this content. So a lot more detail than I can give you today on the semantic layer and why that is so important for context and AI. Last week, we did a, webinar on the future of analytics looking at BI three point o and what that means. We are on the semantic layer today, and we've got three more still on the on the way, in very close proximity the next two weeks, and then one after a little bit further away. And then we have another one already set up for October. So a lot of great content coming to you from your friends at Interworks. We'd love to see you at all of these. And if you do happen to miss, register for them. And then that way, when we do post the recording, you'll have it, and, you can act you can access that at your leisure. So lots of great stuff. We're gonna end the year with a lot of great content. A little bit about Interworks. If there's nothing else that you take about take away from this webinar about who we are, it's it's this. It's we do data strategy solutions and support. So understanding the direction that you want to go, how you might get there, building a plan or a road map, building the things to help you get there, which might be a data platform, implementing governance, data pipelines, analytics, AI solutions, all of that. And then supporting you, once you've gotten there or on the way, to be quite honest. Another way to think about this, if we just double click down from strategy solutions and support, you can see a clearer picture of more things that we do. So from a foundational standpoint, that could be your your infrastructure, your platforms, your applications, your cloud governance, obviously, is critical, particularly in the era of AI. So that would be metadata quality, cataloging, master data management, you name it. There's a whole bunch of stuff that goes under that umbrella. Your data, which could be your data foundation, for instance, your cloud data warehouse, lake house, your data pipelines, all of that stuff there, we can help you across. We love Snowflake, but we can help you in other places too. And obviously, use a lot of different transformation tools, so we can help. Once you get the foundational pieces, then then you can start looking at really adding value or solving problems. And we can do that through analytics, advanced analytics, AI, data science, machine learning, a whole bunch of things, operational reporting, data apps, you name it. And of course, again, that sustain and grow. So we can help support your community. We can do that through championing your community of practice and seeding ideas, gamification, new content, skill pills, best practices, all those sorts of things. We could support the individual by helping you build more skills and capabilities and proficiencies, maybe even approaching and going towards certification on a particular tool. We can do that through traditional classroom training, one to one mentorship, hands on workshops. There's a bunch of ways. And again, if you need any support with the applications or the items in those applications, like a specific data pipeline, for instance, it's critical and you want someone to look after it and and you would prefer not to, we can help you anywhere and in all of these places. We very narrowly focus on data, but we do all of the things. And as a result of us focusing on data and things related to data like analytics, for instance, it's given us a tremendous amount of experience and pedigree and focus over the thirty years of our existence. In October, not too far away, there's two huge milestones. One, Interworks is going to turn thirty years old. We've got a big bash globally that we were practicing and planning and preparing. Yours truly, by the way, will be turning fifty. I don't think I'm as excited about that milestone. Maybe I will not be doing a bash. I'll just be hiding away and pretending it didn't happen. Some other factors, factoids about Interworks. Over our thirty year specialization and focus on data and analytics and helping folks, we've gotten to work with a lot of really big name companies. Seventy five, maybe even seventy seven of the Fortune one hundred are or have been our customers. And you can see lots and lots of stories about who we've helped, how we've helped them on our website. Obviously, we can't share all of those stories, but we can we share what we can. Also on our website is our industry leading blog on data, somewhere around four million page views per year just on data. That is an crazy amount. We spend a lot of time and energy on it. We put a lot of white papers, videos, customer stories, all sorts of stuff on there Just try to make it as interesting and as helpful as possible. One thing that's great is going to a conference and having a lot of people know me before I get there because, oh, I read your blog. I read your specific blog article. Or I know Interworks through all the great content you guys do. So I would strongly recommend going to interworks dot com, finding the blog, and bookmarking it because there is always good stuff in there. We've gotten to help a lot of customers, as I said, somewhere around four thousand, maybe five thousand by this point, across every industry vertical, all sizes, public sector, private sector, true global enterprise, all the way down to the SMB. It's not really a particular type of customer. It's a type of customer, meaning the relationship that we're looking to build. And we've gotten a lot of awards, and one of those is a Forbes small giant. You can read more about that on just if you want. So let's get started. We got a lot to get through today. Lots and lots of content. I struggled on trimming this one down to something that was a neat forty minutes. So let's just get right in because there's tons and tons of stuff to talk about. The first thing I wanna explain, hopefully reinforce, maybe hopefully you are already believers, but the semantic layer is super important. And if you don't trust me telling you this, well, here's a whole bunch of other people telling you to. Really, really important people, really, really important companies, and all within the last, say, two to three months. The one quote on here that I like the best, the semantic layer has always been essential. Now it it's existential. I think that is probably bang on accurate. I'll explain why if you don't fully understand just how different things are today in terms of building that semantic layer than it was, say, two years ago. But let's start by looking a little bit backwards. We're gonna go all the way back to nineteen ninety one, which is really the the origin story, the superhero origin story of the famed and fabled semantic layer. Business objects created the semantic layer. They did not call it the semantic layer. They, of course, they branded it as the business objects universe. But basically, the whole idea was is we we have people that want to understand business concepts, that want to see these things, but don't wanna learn the technical bits. And literally, some vice president didn't wanna learn SQL. So they built an abstraction layer to make it easier for this person to understand what was actually happening in the data. And if abstraction layer doesn't make sense, we'll get to the definitions later. This is meant to be as approachable as as possible for everybody. So if you're a data person, hopefully, there's some useful stuff in here. If you're not a data person, hopefully you'll find your way along too. So BusinessObjects was bought by SAP in around two thousand and eight. Lots and lots of money. Yay whoever created BusinessObjects. But we've been on this journey of the semantic layer for thirty five plus years, and it is going through a massive revolution today. So from there, we have the nineteen ninety one to the two thousand. So this monolithic era, which is invented by business objects, quickly pilled into by other companies like Cognos and MicroStrategy. And yes, it was easier than, say, SQL, but it was still very proprietary, which meant if you use business objects, other tools probably, most definitely, could not look into that application and understand what was happening there and use that business logic anywhere else. It was very much a a lock in. And that was primarily for two reasons. One, it was to the advantage of business objects to lock you in. And then two, probably limitations in terms of the technology as well. A lot more advancement has gone to making these languages and this logic more extensible. That doesn't mean there's not the incentive from vendors to keep you, inside of their proprietary system. So somewhere around the late 1990s to 2010s, that's when you had that multidimensional era. That's when the whole cube thing happened. And they were revolutionary because they were highly performant. And then back in those days, we're talking still twenty to thirty years ago, the amount of data was faster or there was more growth in data, there was more of data than there was processing power. So that everything we could do, and it's still true today to a large extent, everything we could do to accelerate the ability to process data, to aggregate it, to make sense of it, was super useful, super valuable. And that was the big advantage of cubes. Pre aggregated, highly performant. The challenge was, is that to pre aggregate them, you had to do a lot of highly rigid work to get all of these things lined up. And so if you added something, changed something, it was a ripple effect that a lot of things had to be changed to get all of this to work. So while it was useful, it wasn't the endpoint. Not that there's ever gonna be an endpoint, but it wasn't the direction. It was sort of like, let's go forward here. This looks good. Okay. We should probably take some lessons learned. Take a step back and try something else. And that's where you get to twenty twelve, semantics as code. Looker, and this was pre Google Looker, introduced LookML, which is moving the semantic model out of UI menus and into version controlled code. And again, look at MEL, proprietary, non portable, particularly at that time. But it started making this easier, better, and something closer to what we see today. We then go forward to the 2020s. That's where we start to see this headless or universal era for the semantic layer. This is the rise of tools like DBT. This is where the semantic layer steps away from BI tools entirely, like Looker or Business Objects. And the idea of it being headless is that it can sit directly on top of the warehouses like Snowflake. And it can service any tool by an API. So I don't have to have Looker to use my LookML semantic layer or some complex conversion thing to then get it to sort of spit out so that some other tool can use it. It is agnostic intentionally. It's agnostic by design, and that's part of the value. Because we're going to have a lot of different types of things that want to consume it, which brings me to our latest sort of thing. One of the big consumers of that is going to be AI. And that is true for everything. AI is rewriting the way we think about all technology. There is in fact, a massive effort just on the web to go and redo all the websites and redo everything that's happening that was built for humans to consume it, for agents to consume it. And of course, that's a hundred percent true at the data layer, the semantic layers, everything now. The number one consumer in the next two or three, four, five years are going to be artificial intelligences and not humans. We're going to be the people behind those while the AI does all the work for us. So we've got two primary audiences. We still have to write for human beings so it's easy for them, but we have to take into account AI. We're gonna get back to how all of this translates into what semantic means today. But let's start with some base definitions first. So key architectural concepts. So the very foundation of all of this is business logic. What is a customer? What is revenue? These are decisions human beings have to make about how your business operates. What is a transaction? What is an event? What is a thing? How does that relate to something else? That's your business logic. And you can translate that into code, SQL, Python, whatever. Now, that's very technical and technical people can look at that and kind of figure it out. Mainly everyone else cannot. And this is what the idea of an abstraction layer is. What an abstraction layer is, and they can be all sorts of varieties, there's not specific to this context of data only. But an abstraction layer takes something that is complex and then makes it more simple. It abstracts away the complexity so that it's easier to understand. So a semantic layer is an abstraction layer. Not all abstraction layers are semantic layers, if that helps. So a semantic layer. It's an abstraction framework that sits between the raw data and the end user tool. So it's that middle point where it is easier to understand and more simple and efficient to use on the consumption side, meaning Tableau, Power BI, Excel. It gets a little bit more complex. So ontology. This is the conceptual blueprint. Think of it as purely logical and it's how things relate across the business ecosystem. What entities exist and then how they all come together. Super important from a governance perspective when you're thinking about data domains, entities, etcetera, etcetera, but important. Then we have the metric layer. This is something that popped up, if memory serves, from more of a DBT context. But again, great ideas come from everywhere and it very quickly get adapted and adopted. We'll see that more stuff with some stuff that came out of Databricks in this very presentation. But here is where we're thinking about critical business definitions, the mathematical formulas, aggregations, all of that stuff in one centralized place, strictly focusing on numbers, math dimensions, and aggregations. All of the business context around that can sit in the broader semantic layer. The metrics is where we want to actually lock in these formulas and definitions for broad reuse. And the name of the game here is broad reuse. I'll throw this out here. It's not specifically tied to semantic, but it is an acronym that you'll hear quite a bit, which is RBAC, Role Based Access Controls. Meaning who gets to access based off of what permissions they have, what role they have. And the more you simplify this, so I am an analyst, I get to see what analysts see, not what a special Robert usage profile gets. Look, the standardized, easier, more predictable all of this access and security and privacy is going to be. And as you have more and more data, and more and more of that data is going to be sensitive PII or regulated or compliance related information, the more robust, more thoughtful you need to be in terms of how you police access to your information. Super important. I think that's all of my definitions. Let's hop over here. This is another thing that's come up. There's a lot of different ways to slice this by multiple axes or multiple dimensions. This is one that's come up recently, which is the open semantic interchange. Snowflake, DBT, a whole bunch of people have gotten together said, hey, let's all agree on some basic definitions here. And so if you look at the data layers, layer seven all the way down, and its corresponding traditional OSI equivalent, you can see the semantic layer is now nicely settled into layer six, which was that presentation layer, or the translation layer. This is where your business logic and metrics live. So there's a lot of different ways the industry is trying to formalize this, standardize it, define it so that it's easier, more approachable, and more consistent. Okay. I mentioned Databricks before. Databricks introduced a couple of things. For instance, the lake house that was originally a Databricks concept, but made sense. So we started using it everywhere. Another one that was introduced from Databricks is this idea of medallion architecture. And regardless of what tool you're using, you're probably going to it's very, very probable that you'll end up using these terminologies because they're quite useful. And there's three layers here. There's bronze and, think of this medallion like the Olympics. So you've got bronze, silver, gold, and you might have other metals on top of that depending on what's useful for you. But they each represent a different layer of complexity or think of it as a manufacturing process. And each of them represents a stage in polishing and finishing data through DevOps cycle. Bronze is where we got through the APIs, we dumped all this raw data into our data warehouse. Unstructured, unvalidated, completely abstracted from the business. They don't need to see this raw data. Our intent is just to get it all in one place. We've got through the APIs. We got through other data connectors we had to. We then get to the next layer, which is silver. Cleansed, deduplicated, conformed. We've started to build some relationships. Now sometimes you could point your semantic layer directly here and then build things off of it. Just depends. What are your goals? What does the silver actually look like for you? How complex is your data? But people have and can use the silver layer if they want to go do stuff. There's going to be more work that they have to do, but they don't have to build a finished data product at gold, as you can see here, to do something with it. So I don't I don't wanna give you the impression that it's gold or bust. But gold is obviously the highest level of data product that you can build here. You've got your star schemas, facts, dimensional tables. Again, still just rows and columns. The logic is still exposed, but it is far more prepared, far closer to something that a business user could leverage and take advantage of. Now, these terminologies are useful, but they're not consistent everywhere. If you were to go and sit in an architecture meeting in company a, let's say they're in public sector, and then you were gonna send another architecture meeting over at company B, and let's say they're in supply chain, or they're a digital native, or they're in mining, or whatever. Even if they were all using medallion architecture, they might use them in slightly different ways. Because again, it's not written in stone. These are ideas and concepts and people will mold them to what's more useful for them. And that's within the context of the way they do business, way they think about their architecture, etcetera. I say that because I've got the semantic layer sitting here, but other people might put the semantic layer in different places. There may be a platinum first, or they may pull this into semantic sitting right next to context. They might put context inside of semantic. It does vary. But this virtualizes the gold layer, defining all the relationships, locking in formulas. And ideally, this is where you want the AI to consume from because it's the easiest thing for AI to look at and understand. If I see something that's obviously not right or a definition that was not accurately done or missing something, I can, as human being, I can second guess, I can question, I can go explore and validate what I'm looking for. AI is not as good as that unless you tell it to. What AI is good at is looking at a hundred million numbers and saying, that average doesn't look like I can't see that difference. So we have different things that we're looking for. And so if we build the business logic, all the rules into the semantic layer, that's the kryptonite for AI and that's what it needs. So that's why you want to put all of that polish from bronze, silver, gold, and into semantic so that AI can benefit from it. So if you think about the benefit of the semantic layer more broadly, think of it from a bottom up now. You have to have a well built semantic layer to power a lot of things. BI, all of your dashboarding, your reporting, you need business logic. Now, in the old days, you could build your semantic layer as you go. So here I am in the data warehouse, I've got, this table is really good. I then go to my Tableau server and I've got a published data extract there where we've got a lot of work inside of that. So we'll grab logic out of there. And then I get down to the workbook itself, maybe to the worksheet, and I start to type my own little calculations. Maybe I bring in my own Excel file or SharePoint doc, and I get the rest of the way there. At the end of this, I have created an automated dashboard that probably gives accurate answers, if not completely accurate answers. That's not a semantic layer. That is distributed scattered everywhere. It works for the sort of piecemeal breadcrumb approach for BI reports, but it doesn't work anywhere else, particularly with AI. So if you build the semantic layer as a consolidated load bearing structure, everything here will benefit. Even the BI, where you could sort of piecemeal your way through it because you're spending less time and energy to have to do that. And it's going to standardize your definition so there's not chance of forking or fracturing a business logic. So AI, ML, data apps, your data products that you're building, exports and shares, for instance, in you know, you could build that marketplace across the Snowflake sharing. All of these things benefit from. So if you wanted to productize your data and sell it, all of these things benefit from having that beautiful semantic layer. Now, it's different than it was in the past because as I said, the semantic layer as the load bearing structure, everything has to be here for this to work. That's the big shift. In the past you could piece it together in a dashboard, but it's not just human beings and dashboards that's consuming data. It's all of these things. So that's the real paradigm shift. So let's look at an example. The initial metric debt requires time and business alignment to reconcile years of fragmented definitions across different platforms. So here we have lots of sources of data. They're different varieties. Maybe some of them HR, some of them are finance data, some of them are sales data, some of them are from the CRM or ERP or or blah blah blah blah blah. We have to get that data into our data warehouse so we can reconcile it, you know, bronze it, silver it, gold it. We'll do that through our ELT process with our data engineers steadfastly working away. And then we get something with all the context, all the business logic, all of our semantics metrics, all of that stuff defined into a beautiful semantic layer, which is then is consumed by a dashboard built by an analyst. And then users consume that, and they get insights. Now we just talked about BI three point o in our very last webinar last week, and we talked about how this part of the paradigm is shifting and how that changes everything. So it's worth having a look in terms of what smarter, more capable AI augmented analytics are going to do to the entire stack down the data life cycle. Now, the way this normally works is, and I've seen this play out over hundreds of different organizations, data engineers are generally centralized. You might have some specialists that are popped into some very data progressive organizations, your finance, your marketing, whatever, product, whatever. But generally speaking, data engineers are a centralized resource. And that makes sense. A centralized data team, because you would centralize the selection of the data platform. So the people that look after and administer the platform were there. The people that built the data platform, the data pipelines rather, are all in a similar place. Then everybody can benefit. It makes a lot of sense, but you don't get a lot of investment in data because they are centralized. If you've got ten engineers in there, you're like, well, we don't need twenty, do we? The challenge is that while data and platforms is often centralized, analytics, particularly in today's world, with Power BI and with Tableau and Looker, and the newer tools like Sigma, etcetera. Oftentimes analytics is federated because we don't wanna wait for IT. We don't wanna wait for engineers to build us reports, or dashboards, or help us do our analysis. The business can do that. These tools are easy enough that I don't have to be a classically trained data engineer or statistician to do it nowadays. So often what we'll see in the business units and the teams is a lot of analysts. And you can imagine those analysts are going to be moving faster than the engineers can, which then creates this, well, I don't have data over here on the left hand side. So I'll start doing stuff on the right hand side, which is what I can control and I can get my analysis out first. And the logic might be in the workbook, it might be in Excel sheet, it might be in an extract that I pulled directly from the source system, which is maybe this rock pile over here. Because again, I have to wait for the data dev ops cycle to get to what I need. And on average, data estates grow by forty percent a year, or that means they're doubling every two years. So there's always a backlog of stuff for engineers to get to. And so what that means is, is if the logic starts to exist in other places, that means you don't have the semantic layer anymore. That means your semantic layer is incomplete, inconsistent, inaccurate. And that then becomes well, now the dashboard is the load bearing structure. That's the only place that represents the single source truth because we have this imbalance of analysts and business people needing information, data insights, innovation faster than what the data DevOps cycle could keep up with. And of course, we proliferate more data. We always have more data. And that just creates that backlog on that side. And maybe, I've seen this happen quite a bit, data engineers are like, we can't get to silver. We certainly don't have time to get to gold. So we'll just get everything staged into bronze. And they're doing everything they want in the analytics layer anyway, so we'll stop there and just try to keep our head above water, which becomes a compounding problem, which you can imagine everything starts to facepalm. This gets even so this wasn't great. We could still get dashboards with answers, but this gets even worse when you start to think that the primary consumer of all of this is not gonna be the business user anymore. It's not gonna be the analyst anymore. It's gonna be an agent. And the agent is going to look at this and then give the answers directly to the business user. I, as a business user, would say, what did our revenue look like for this region, for this business unit, for this time period? And the AI is gonna go and come back and give me the answer. But if it's a fractured, fragmented semantic layer with our business logic everywhere, the AI has just as much of a chance of getting it right as me throwing a dart at a board. This is where a lot of organizations find themselves today. They made BI two point zero work, the world of ten thousand dashboards work through duct tape and a lot of extra work at the analytics layer. And that's the problem because the future of data, it does not look like more of that. Semantic is the load bearing structure. So we have got to get all of these things require the semantic layer. We've got to get the semantic layer consolidated and curated. It's absolutely mission critical. We could fake our way through with BI two point zero. We can't do that today. So to help, I'm gonna give you ten critical things to keep in mind. And these range kind of a lot of different places. Just kind of a grab bag of things I thought were important. So number one, centralize your business logic. A single source of truth. So you cannot have isolated BI reports or PowerPoints or Excel or that data source, or we have multiple warehouses and sometimes they overlap, sometimes they don't. You get a benefit at the economy of scale. You get the benefit from everything being together. It's like going and trying to buy a fleet of cars for your business. And we're gonna buy a car from every dealer in the state, or we're gonna buy ten thousand cars from one. There's a value in doing it together, and it is a compounding value. So centralize your business logic as much as you can. It will illuminate code duplication. It will standardize. It will create consistent business definitions. It will highlight the fault lines where the business does not agree on these definitions. For instance, what is revenue to finance and what is revenue to sales? One is about reporting to the street and and and showing numbers and profitability. The other one is paying salespeople. They don't care about the street. They wanna make their number and they wanna get their variable pay. Different numbers, but the same thing. You've got to come up with human beings making decisions. So if you centralize it, you get the benefit of talking about this and standardizing it and benefiting it, or rather defining it once and using it everywhere. Super important. Number two, avoid the silo. Have seen this multiple times and it actually happens probably more than you might think. And doing a migration from one data platform to another is hard. It takes time. It takes years, quite honestly. So if you're going from Teradata to Snowflake or Redshift to GCP or whatever, It is not a quarter effort. And I think because some people think it can happen quickly, they may not have the organizational persistence and focus to finish. And there might be some folks that are really driving this. And so yes, we've got this particular use case over to this new platform. And now we're going to focus on driving value because gosh golly, we can't wait any longer on trying to get these two things turned over. We got to start driving value. You can't just keep doing this migration. And that's a lack of sort of understanding of how much effort it takes. And what happens is you end up with two data warehouses. This one's over here, tracking new numbers, doing old stuff. And this one's over here also with new numbers, but tracking new stuff. And you end up with both of them running concurrently. And you'd be surprised a lot of times, oh, look, another a third one's popped up. Or we use SAP over here. We use Salesforce over here. And then we've thrown Snowflake over here. Again, numbers everywhere. And people don't know what to turn to. So that means when you have data that is trapped, you have silos, which means inevitably there's fracturing and forking in terms of business rules and definitions. Number three, foundations for governance. There are certain things that you can do from the bottom up, from the community upwards, like ideas, innovation, driving cool stuff that you might think of in terms of dashboards and reports, new products. That really works well at the coal phase. What does not work well from the bottom up is things like strategy or governance. They have to start from the top down because governance in particular is a mandate on control, on standardization. It's not creativity, it is the opposite. It is conformity. And so you have to have important people with big titles and with a big stick saying, stop doing that and do it this way. The semantic layer is the foundation for governance because it can sit in a centralized repository, right where it has come out of abstraction so that the users can see it. So there's this very narrow limited use, limited access machinery behind it, bronze, silver, gold, and then boom, big debut, everyone can see it. So if we put governance there versus at other places, in other tools, or other layers, you then have a problem. And when we're talking about data governance, talking about access. So row and column level security, we're talking about classification, metadata, data discovery, all of those things work well when you have a consolidated semantic layer. We've already talked about semantic layer and AI. But honestly, if you do not have a semantic layer, you can still do AI. It's just gonna have a couple problems. One, it's gonna take a lot more time to build your solution for production. Two, you are likely to get a great idea that cannot get out of your prototype because your data isn't good enough to validate a great idea. Or three, the model might work. Your data might be good enough, but it's going to be unpredictable over time because you have no governance to train the model and keep make sure that you've got the right data in there. It's a whole accounting principle, garbage in garbage out. LLMs, generative AI agents, they they they're going to report back an answer whether they have a good one or not. So a well curated semantic layer, particularly with that unstructured data, they can all come together and create a full picture. Super, super important. It is the rocket fuel for AI solutions. Productivity gains. So, one of the easiest things you can do to validate this, when you're when you're sort of trying to sell a big enterprise, initiative to your leadership, you can sell value. AI is so amazing. We do this. It will be a differentiator for us. It is a value creator. We'll do things better. You can also sell the mitigation of risk, or or cost savings. If we have a semantic layer you know how all you know we have a hundred analysts? They take twice as long to build stuff because they're building it from scratch. It is, artisanal in their ability to build data products. It is handcrafted data products that they're having to do in Tableau. If we built it in the in a data factory, well, then they could just use our products and that would accelerate the time for them to generate insights by twenty, thirty, forty, fifty percent. Cost savings. And depending on what your organization wants to do, you might redeploy those people into other areas. We don't need fifty dashboards anymore. How about thirty and then we have folks over here building more data products, or we have folks over there building data apps or building agents, adding value, new things. Productivity gains is a great way to sort of frame one of the key values of the semantic layer, easier to build stuff with better data products. Multi tool interoperability. One layer to serve the entire data ecosystem. Again, if you've got your semantic layer in a tool like say DBT or Snowflake or some other agnostic extensible tool that you can just point another tool out, it means you don't have to build it again and again. You don't have to do the extra work to translate this proprietary code from say Salesforce into something that another tool can use or hope and wait that they do it for you when they get around to it. And wow, aren't they super motivated to give you the ability to take the business logic out of their platform? I'm picking on Salesforce. That's every company. Every company wants to maintain customer retention. So it's kind of not in their best interest. New paradigms, embracing trends to accelerate. This idea of, a a middle ground forming between data engineers and data analysts is not the newest idea in the world. It's probably been around for at least eighteen months to two years, you know, at some level of, some tipping of the scale to where it's worth taking notice. But it is coming to, I think, a significant, headway. And what this is, is this. The ability the the need and ability for people to write to to do data transformation, to build data products is one, easier and more efficient to do today because there are tools helping you do this. There is AI helping you write code. There is AI helping you navigate APIs. So the ability, for people coming from the analyst side, from the business side into this data, transformation effort is easier than before. The other part of this, and I already mentioned this before, is the need for data products is substantial. It's the number one thing that you can do to unlock yourself. More data products, more gold layer stuff, more semantic layer stuff, more things people could use to use AI, to use dashboards, to use advanced analytics, all of those things become critical. So what we want is we try to empower and encourage as many people as possible writing good data products. So what does this look like? If you remember our little example before, we had the data engineer looking at the source system, doing the ETL, and then building a nice semantic layer, I. E. A data product. So the analyst could then do some analysis with their dashboard. Well, what's changing is the analyst now has the ability to go a bit further up the data stack. And depending on the analyst, if they know SQL and they know Python, they might even go all the way back into writing gold or some derivation of gold, like a gold specific or platinum layer that's a bit more use case specific versus this broad use that has to be, you know, spend more time and energy building up. The more you can retrain your analysts to not think about just dashboards, but build data products, which allows them to build data apps, allows them to build generative AI embedded into dashboards, as well as other data products and AI agents themselves, the more useful all of this is going to be, how much more accelerated you're going to be in getting value out of all of this. It is very much a new trend. And as you guys are thinking about what you're going to do with your teams, your team of technical people, BI people, data people, you should really be thinking about the analytics engineer as a fulcrum that you can put into your business units. That means they understand the business context and business domain, but they've got the data chops to actually write up. Think of it like a data mesh, where these are the people that might be looking after these contracts. Maybe not all the way down the stack, but they're the ones that are contributing to the business logic from their unit, as well as being able to to build and explore the the the dashboards, the reports, the analytics. Let's go to the next one. Number eight, people define truth. The semantic layer is not software. It's not just software. It is not just AI agents. It is not just generative AI spitting out thousands of different metadata tags. Ultimately, this comes down to people. You cannot replace people in this process. Even when you look at the technical roles, your data architects, the people that actually do the modeling, they'll probably be able to do a lot more things and juggle a lot more plates, but you can't take those people out even in the five year, seven year window. Because human beings have to make these decisions on what's important to us. So ownership and consensus. And this comes back to a couple of different facts. One, it is human beings need to make decisions, which means that is how we intersect some of these business rules with governance. Let's decide our business rules, let's decide our business definitions, and then lock those into the centralized repository of our business logic in our warehouse, in our semantic layer, so then everyone can benefit. And where we need to, let's have those battles, let's draw those compromises and put the asterisk where we need to, but let's get these things going. People define truth, you need the ownership consensus. That's the idea of data stewardship, data ownership, technical ownership, a lot of different ways that you can sort of slice and dice this from a governance perspective. Cost optimization. Here's another one that a lot of the important people with c levels and their titles like. Without a semantic layer, pulling together all these definitions and then having it nicely curated and hot on the stove, ready for you to use, your analyst might actually go build their own query separately, independently, redundantly, multiple times. You might have people all over the organization, and I've seen this, running queries that basically do the same thing. It's bad enough that they might be slightly Oh no, that number is different from that number. And it's going to show up in our executives report. Marketing's going to have one number, finance is going to have a number, but they're also just replicating cost. And everyone has this cold shock nightmare fear, which they should of waking up like, oh my gosh, we just spent four hundred thousand dollars on ChatGPT because we weren't watching tokens and credits. So cost optimization, particularly on anything that's consumption based like your data warehouse, like Databricks, Snowflake, whatever, like your AI tools, like Claude, whatever, you want to be intentional about how you spend your compute. You wanna be intentional on how you spend your tokens, and the semantic layer will help. Pre aggregating data, optimizing paths, having the architects do the really cool stuff behind the scenes will also boost query performance. It'll also lower those costs, and it'll make everything more predictable, and every finance person loves that. Tell me how much I'm going to spend this month. Great. That's exactly how much we spent. I don't like surprises. The path forward. A data strategy without a semantic layer is just a connection of collection of disconnected pipelines. I've said this again a couple times. The way that you need to go forward is go look at where your business logic is for some of your most critical reports. So if you say, hey, this is the this is our executive summary report. Let's say it's a Tableau report with six different dashboards, and it looks at different aspects of my business. Do a little exercise with your data people and say, hey, I want you to tell me how much of this is written in this workbook. Like, many, calculations are there in this workbook? Okay. Now tell me how much custom SQL is in this workbook? And is there anything published to the analytics server that it's using? And is there any third party file like an Excel or anything that's sort of added to this? Trace it all the way back up your data life cycle and say, does it actually sit and and and document it? That's how you'll get to understand where your actual semantic layer is. If it just so happens that everything is nice and neat running off of a live query from Snowflake, right. Now you can start to think about other business units. But until you do that exercise of understanding where the fragmentation is, you can't fix it. And you have to fix it so that you can benefit. I'd also have a look at, the marketplace in terms of other tools. So obviously Snowflake does a lot of cool stuff. Databricks has, cool stuff in there as well. We love Snowflake. There's a ton of amazing stuff that they released this year. There's more amazing stuff. There's other tools that you can put with it to compliment it like a DBT, for instance. If DBT is a little bit too technical, you could also compliment that with something that is a bit more gooey face, like a coalesce or something like that, where people that aren't technical can do this from very much a visual workflow. But all of these things together will help you. And again, it's very dependent on what your needs are, what you're trying to accomplish in terms of the actual blend. And someone like Interworks can help you. If you say we have fifty ******** SQL Python developers, that might be one solution stack. That might be one operating model. That might be one shared authorship ownership data DevOps model. If you tell me that you've got one person that knows how to do everything and everything is capitalized into their head, but we need to accelerate. And we've got non technical people that are willing in other departments, that's a completely different solution. And they're both very realistic scenarios. But we can help in terms of figuring out how best sort of array you to get you value today and then set you up for long term success. Hopefully, found this useful. There are ways we can help. I'll give you an example, but quite frankly, the the the number one thing that you probably should be thinking about is, hey, Interworks, I need to talk out where we are right now. And I need to understand what is the things I need to do in the next three to six months. And then the next one to two years. The next three to six months, I would say these are very specific tactical things that you should be doing. And I don't know what they are. I have to talk to you to understand what your needs are. But I'll you some examples. Let's go figure out what your primary use cases are, your high value use cases, and map where that semantic layer is and figure out what it would cost to centralize it. And then we'd map that up against the ROI and use cases we'd solve and see if it made sense. The longer term stuff is more directional. And the reason I say that it's not because I'm trying to be wishy washy. It's that because the speed of change and innovation means it's not smart to put big concrete bets on something in two years or three years. There are predictions that AI could accelerate the bronze layer development from an engineering standpoint by twenty percent, or maybe even up to eighty percent. That's a variance. Every time I turn around every quarter, it seems like Claude's ahead or this model's ahead or that model's ahead or this model's ahead. So I would say let's, let's focus on what we can in the next, say, six to twelve months. And there's some foundational stuff that readiness type stuff, which is governance, strategy, data, and people getting all of those things lined up, looking for opportunities to do prototype AI solutions on the dataset that is the nicest, most polished, most curated, most centralized, learn, add value, and hedge. And then there will be more developments, more value, more accessible as more and more of the data becomes, or not more of the data, more of these solutions become, productionized, proven, tested, practical. If you are an early adopter and you have the appetite for that, it's a completely different conversation. Most folks, particularly in today's economy, we need to have results. We need to be very, very mindful of the money we spend, and anything that we can play towards efficiency is going to play well to our leadership. That's generally where people wanna talk. We can do that. We can help. If you are looking for new tools or strategy or whatever, there's a lot of things we can have. I've got up here a spec and select process. And if you're thinking about things to spec and select, I would say it's probably governance related. It is your data platform. You'd want a cloud based data platform like Snowflake. Snowflake obviously has a ton of amazing stuff, but it is growing horizontally. So it can do data science and data cataloging and, containerized LLMs and AI models and natural language to code, all kinds of amazing stuff. Let's have a look at it and see how it'd be useful it'd be for you. Also, could be looking at new BI tools that integrate AI, that sit natively right on top of all the cool stuff Snowflake can do. There's a lot of different things that you can look to really accelerate you, and we can help you do all of that. If you wanna talk to us, obviously, you can just go on to our Interworks website and contact us, or you can scan this QR code that I've got in front of you. It'll send an email, I think, to our sales folks. We would love to chat with you. Love to see how we can help. Certainly would love to see you show up on our webinar next week and the week after and the one later on September and the one that's in October. I think we might do one in November too. Hopefully, found this useful. If you have any questions, we have about five minutes left. I see Rafael. Thanks, Raf. It's a pleasure. Thank you for the nice words. If you do have any questions, I'm happy to answer whatever you throw in there with whatever time we have remaining. Otherwise, if you are good and you have everything you need, we will shoot out the recording in a couple of days and you can share it with your friends, whatever you need to do. We'd certainly love to see you guys again. I'll hang out here if you guys have any questions, though. Alrighty. It looks like we're all good. Thank you, guys. Thanks for joining. I'll see you next time.

In this Webinar, Rob Curtis reviewed the history of the semantic layer, its practical use and the tools you can use to implement it in your own data.

InterWorks uses cookies to allow us to better understand how the site is used. By continuing to use this site, you consent to this policy. Review Policy OK

×

Interworks GmbH
Ratinger Straße 9
40213 Düsseldorf
Germany
Geschäftsführer: Mel Stephenson

Kontaktaufnahme: markus@interworks.eu
Telefon: +49 (0)211 5408 5301

Amtsgericht Düsseldorf HRB 79752
UstldNr: DE 313 353 072

×

Love our blog? You should see our emails. Sign up for our newsletter!