<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Newsletter &#8211; Prologika</title>
	<atom:link href="https://prologika.com/category/newsletter/feed/" rel="self" type="application/rss+xml" />
	<link>https://prologika.com</link>
	<description>Business Intelligence Consulting and Training in Atlanta</description>
	<lastBuildDate>Sat, 19 Sep 2026 18:48:44 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>
	<item>
		<title>Prologika Newsletter Fall 2026</title>
		<link>https://prologika.com/prologika-newsletter-fall-2026/</link>
					<comments>https://prologika.com/prologika-newsletter-fall-2026/#respond</comments>
		
		<dc:creator><![CDATA[Prologika - Teo Lachev]]></dc:creator>
		<pubDate>Sat, 19 Sep 2026 18:48:44 +0000</pubDate>
				<category><![CDATA[Newsletter]]></category>
		<category><![CDATA[Excel]]></category>
		<category><![CDATA[Fabric]]></category>
		<category><![CDATA[Semantic Model]]></category>
		<guid isPermaLink="false">https://prologika.com/?p=9715</guid>

					<description><![CDATA[Ask a business user where they want to work with the data, and Excel would probably top the list. Regardless of how hard we try to lure users away from [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><img decoding="async" loading="lazy" class="alignnone size-full wp-image-9657" src="https://prologika.com/wp-content/uploads/2026/08/word-image-9656-1.png" alt="" width="1408" height="768" srcset="https://prologika.com/wp-content/uploads/2026/08/word-image-9656-1.png 1408w, https://prologika.com/wp-content/uploads/2026/08/word-image-9656-1-300x164.png 300w, https://prologika.com/wp-content/uploads/2026/08/word-image-9656-1-1030x562.png 1030w, https://prologika.com/wp-content/uploads/2026/08/word-image-9656-1-768x419.png 768w, https://prologika.com/wp-content/uploads/2026/08/word-image-9656-1-705x385.png 705w, https://prologika.com/wp-content/uploads/2026/08/word-image-9656-1-450x245.png 450w" sizes="auto, (max-width: 1408px) 100vw, 1408px" /></p>
<p>Ask a business user where they want to work with the data, and Excel would probably top the list. Regardless of how hard we try to lure users away from it, Excel remains the endpoint of self-service BI for many organizations. I have clients who have built custom apps or purchased Excel add-ins simply to get data into Excel—or manipulate it once it gets there.</p>
<p>So rather than fighting Excel, let&#8217;s look at the options Microsoft provides for bringing governed Power BI data into it. In particular, I&#8217;m interested in getting data from Fabric semantic models into Excel.</p>
<p>Many things have changed over the years in the data analytics space, but semantic models have stood the test of time. The technology has changed—from multidimensional OLAP cubes to tabular models and now Power BI and Fabric semantic models—but the basic idea has remained remarkably consistent: put business logic and governed data in a centralized model, then let different tools consume it.</p>
<p>Here are the main options Microsoft provides today to make data from semantic models available in Excel:</p>
<table>
<tbody>
<tr>
<td><strong>Option</strong></td>
<td><strong>Pros</strong></td>
<td><strong>Cons</strong></td>
</tr>
<tr>
<td><strong>Export from published Power BI or paginated reports</strong></td>
<td>Power BI collaboration features; sharing; subscriptions; dynamic subscriptions; familiar report experience</td>
<td>Export limitations; additional effort to create reports just to get data into Excel</td>
</tr>
<tr>
<td><strong>Excel PivotTables and PivotCharts</strong></td>
<td>Familiar reporting experience; live connection to semantic models; customizable drillthrough; ability to share Excel reports in Power BI Service and generate them online; OLAP formulas</td>
<td>Outdated report experience; MDX query interface; rigid report layout; Excel workbook must be shared; must open report from OneDrive/SharePoint to interact</td>
</tr>
<tr>
<td><strong>Power BI connected tables</strong></td>
<td>Data is delivered directly into an Excel table; designed specifically for Power BI semantic models; good UI for selecting fields and filters</td>
<td>Less flexible than a full report; users don&#8217;t get the full Power BI report experience</td>
</tr>
<tr>
<td><strong>Power Query</strong></td>
<td>Transform data; mash up multiple data sources; reusable queries</td>
<td>Suboptimal MDX queries, additional complexity; connector limitations; licensing/connection considerations</td>
</tr>
<tr>
<td><strong>Excel Copilot</strong></td>
<td>Natural-language instructions; can retrieve and augment data; can transform data</td>
<td>Still evolving; can be slow; not yet ideal for creating a reusable library of data-extraction definitions</td>
</tr>
</tbody>
</table>
<p><strong>Export from published reports</strong></p>
<p>This is probably the most obvious option. A Power BI or paginated report can serve as the interface for users, with Excel being the ultimate destination. This approach can be particularly attractive when the organization wants to take advantage of everything Power BI has to offer: report sharing, subscriptions, dynamic subscriptions, annotations, and other collaboration features.</p>
<p>For example, a report can be designed specifically for a group of users and distributed through a subscription. Paginated reports provide even more flexibility when the objective is to produce formatted Excel or CSV files on a recurring basis.</p>
<p>The problem is that Power BI report exports have limitations. Depending on the report and export method, there are limits on the amount of data that can be exported. Paginated reports are considerably more flexible for large-scale tabular output. There is also an architectural question: do we really want to build a Power BI report whose primary purpose is to get data into Excel? If the report exists only because Excel is the final destination, we&#8217;re arguably using the report layer as an unnecessary intermediary.</p>
<p>The better architecture is often: <strong>Semantic model → Excel </strong>rather than: <strong>Semantic model → Power BI report → Excel</strong></p>
<p><strong>Excel PivotTables and PivotCharts</strong></p>
<p>Excel has supported PivotTables since the 1990s and OLAP PivotTables for more than 25 years. This is hardly new technology. And that&#8217;s part of the appeal. The experience is familiar to generations of Excel users. A user can connect to a Power BI semantic model, select fields, slice and dice the data, and create PivotTables and PivotCharts without building a Power BI report.</p>
<p>The problem is that the fundamental PivotTable interaction model has changed surprisingly little. Slicers and timelines improved the experience, but the overall model still feels much closer to the Excel/OLAP world of the early 2000s than to today&#8217;s Power BI experience. As such, it generates suboptimal MDX queries although not as bad as Power Query in import mode.</p>
<p>For an Excel user, however, there is one particularly interesting capability: customizable drillthrough. When a user double-clicks a PivotTable cell, the semantic model can provide a detailed table of the underlying records. A semantic-model developer can control this behavior using the measure&#8217;s <strong>Detail Rows Expression</strong>, which can return a table-producing DAX expression. This is a powerful capability when the semantic model has been designed with Excel users in mind.</p>
<p>Another useful feature could be converting pivots into OLAP formulas. This gives finance professionals the precise, cell-by-cell control they need to build tailored financial reports.</p>
<blockquote><p>There is a catch, though. If you are developing a semantic model that must support both Power BI and Excel users, you need to design for the <strong>least common denominator</strong>. Excel&#8217;s live-connection experience does not expose all the capabilities available in Power BI reports. For example, there are differences around field parameters, some filtering scenarios, custom visuals, and metadata behavior. Even seemingly small modeling decisions can affect how the semantic model appears in the Excel Field List.</p></blockquote>
<p>In other words, a semantic model that works beautifully in Power BI isn&#8217;t necessarily going to provide the same experience in Excel.</p>
<p><strong>Power BI connected tables</strong></p>
<p>This is the option that could have the most potential if it wasn’t another half-baked Excel reporting feature. Microsoft introduced connected tables in 2023, but the experience has evolved considerably since then. Excel can now discover Power BI semantic models directly, and users can choose to insert either a PivotTable or a Table from the semantic model.</p>
<p>The Table option is particularly important because it delivers the data in the format most business users actually want: an Excel table. The user can select the fields they need and apply filters through a dedicated interface, rather than having to write DAX or build a Power BI report first. This is a significant improvement over the traditional export experience.</p>
<p>For years, the missing piece in the Microsoft BI stack was obvious: Excel users wanted the flexibility of a PivotTable&#8217;s field-selection interface but wanted the result as a regular Excel table. Connected Tables finally moves in that direction. It is also a much better architectural model: <strong>Fabric semantic model → Excel table. </strong>There is no Power BI report in the middle.</p>
<p>On the downside, the user interface is lost once data is exported. Consequently, the user must change the underlying DAX query (a client immediately dismissed this option after learning about this) if they want to make changes, such as adding additional fields. This could be a great option if there is a way to bring back that interface. Even better, let the user pick fields from Field List as they can with pivots but export to a table. What could be simpler?</p>
<blockquote><p>I&#8217;ve developed the <a href="https://github.com/thracian2015/SemanticTable">Semantic Table Excel add-in</a> to fill in the gaps left by Microsoft. It supports both existing connected tables (created using Microsoft’s Insert Table) and tables built from scratch. It provides a familiar PivotTable-like Field List directly within the table context. And supports both Deferred Mode (apply updates manually to prevent constant re-querying) and Interactive Mode(live field updates)</p></blockquote>
<p><strong>Power Query</strong></p>
<p>Another option is hiding in Excel&#8217;s <strong>Get Data</strong> experience. Power Query can connect to Power BI semantic models and provides considerably more flexibility than a simple connected table. Once the data is retrieved, users can transform it and even mash it up with other sources. For example, a user could combine:</p>
<p>Power BI semantic model + Excel file + another data source → Power Query → Excel table</p>
<p>It also makes Power Query a compelling option when the requirement isn&#8217;t simply &#8220;show me this data,&#8221; but rather: &#8220;Get this data, transform it in these ways, combine it with my other data, and give me a reusable result.&#8221;</p>
<p>The biggest issue we faced with this option was performance. For reason unknown, in import mode the Analysis Services connector generates horrible <a href="https://learn.microsoft.com/en-us/powerquery-m/analysisservices-database">MDX queries</a>, with nested CROSSJOINs, one for each imported field.  This was another showstopper for us. In addition, Power Query is a data-transformation tool, not a simple business-user reporting interface. Users must understand queries, transformations, refresh, data sources, and credentials. Excel&#8217;s Power Query implementation also doesn&#8217;t expose every connector and capability available elsewhere in the Microsoft data platform.</p>
<p>So, while Power Query is powerful, it isn&#8217;t necessarily the best answer for a business user who simply wants to select a few fields from a semantic model.</p>
<p><strong>Excel Copilot</strong></p>
<p>And, of course, there is AI. I&#8217;ve personally found Excel Copilot very useful for tasks such as generating test data, transforming data, and working with existing spreadsheets. Microsoft is now taking this a step further by integrating Power BI data into Copilot in Excel. The new Power BI grounding capability allows Copilot to use governed Power BI data when answering requests in Excel. This is potentially a very different way of interacting with a semantic model.</p>
<p>Instead of teaching a user how to navigate a Field List, write DAX, or configure a Power Query, the user can simply say what they want.</p>
<p>For example: &#8220;Bring me revenue, margin, and customer name for the current fiscal year. Filter to the Southeast region and sort by revenue descending.&#8221;</p>
<p>That&#8217;s exactly the type of task for which natural language makes sense. The technology is still rough around the edges, though. The experience can be slow, and the interaction isn&#8217;t yet ideal for creating a reusable library of data-extraction definitions. A prompt that works for one user isn&#8217;t necessarily a well-defined, portable query definition that another user can reuse with predictable results.</p>
<p>For Copilot to become a serious enterprise data-extraction mechanism, I&#8217;d like to see prompts evolve into something more like shareable, governed query definitions—something a business analyst can create once and distribute to other users.</p>
<p><img decoding="async" loading="lazy" src="https://prologika.com/wp-content/uploads/2017/06/060417_1725_PrologikaNe2.png" alt="" /><br />
Teo Lachev<br />
Prologika, LLC | Making Sense of Data</p>
]]></content:encoded>
					
					<wfw:commentRss>https://prologika.com/prologika-newsletter-fall-2026/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Prologika Newsletter Summer 2026</title>
		<link>https://prologika.com/prologika-newsletter-summer-2026/</link>
					<comments>https://prologika.com/prologika-newsletter-summer-2026/#respond</comments>
		
		<dc:creator><![CDATA[Prologika - Teo Lachev]]></dc:creator>
		<pubDate>Sat, 13 Jun 2026 21:28:06 +0000</pubDate>
				<category><![CDATA[Newsletter]]></category>
		<category><![CDATA[Data Virtualization]]></category>
		<category><![CDATA[Fabric]]></category>
		<category><![CDATA[Lakehouse]]></category>
		<guid isPermaLink="false">https://prologika.com/?p=9636</guid>

					<description><![CDATA[If Microsoft Fabric was the Statue of Liberty, the inscription would be “Give me your data”. Fabric is obsessed with owning the data when it makes sense and when it [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><img decoding="async" loading="lazy" class="wp-image-9535" style="padding: 0px 10px;" src="https://prologika.com/wp-content/uploads/2026/06/word-image-9609-1-1.png" width="215" height="109" align="left" /></p>
<p>If Microsoft Fabric was the Statue of Liberty, the inscription would be “Give me your data”. Fabric is obsessed with owning the data when it makes sense and when it doesn’t. As I wrote <a href="https://prologika.com/first-look-at-fabric-iq-the-good-the-bad-and-the-ugly/">before</a>, this pattern was probably borrowed from Palantir and to align Fabric with the push for “modern” medallion architectures. Or, to establish a permanent dependency on Fabric…</p>
<p>In this newsletter, I make a case that Fabric should support better data virtualization that goes beyond creating shortcuts to files.</p>
<p><strong>Auto-replicating data to Fabric</strong></p>
<p>To satisfy the Fabric data appetite and facilitate data ingestion into OneLake, Fabric offers two primary options that don’t require explicit ETL: mirroring and shortcut transformations.</p>
<ol>
<li>Mirroring targets a growing number of relational and non-relational database engines. Although described as “easy-to-use”, mirroring could prove challenging to set up in real life. For example, in one case, the client simply refused to set up mirroring from Google BigQuery because of the requirement to grant excessive permissions. In another case, we are still trying to figure out why mirroring doesn’t work from Azure SQL MI configured on private network. Not to mention that mirroring even from Microsoft SQL SKUs has limitations, such as historical temporal tables can’t be mirrored.</li>
<li>This leaves with the second option: shortcut transformations. They target a <a href="https://learn.microsoft.com/en-us/fabric/onelake/shortcuts/transformations">subset of file formats</a> (not databases). Like mirroring, Fabric polls periodically the shortcut target folder and synchronizes the data in OneLake Delta tables. These transformations could be useful to provide convenient access to this data from Fabric workloads, such as to access reference data a business user maintains in an Excel spreadsheet in a Fabric Data Warehouse. On the downside, data must be exported and saved as files.</li>
</ol>
<p><strong>OneLake Shortcuts</strong></p>
<p>Yet, many scenarios could be addressed by simply accessing the data where it is. In other words, by data virtualization. As it stands, Fabric has limited file-based data virtualization capabilities with <a href="https://learn.microsoft.com/en-us/fabric/onelake/onelake-shortcuts">OneLake shortcuts</a>. OneLake shortcuts shouldn’t be confused with the shortcut transformations mentioned before. OneLake shortcuts are read-only pointers to external files. These shortcuts are typically listed in the unmanaged section of OneLake (the Files folder). OneLake shortcuts don’t import the data in Delta tables. How are they useful then? The main thought is to conveniently share data between teams, workspaces, or domains, workloads, without moving it.</p>
<blockquote><p>As an exception to the rule, if the OneLake shortcut points to a Delta table, such as OneLake or elsewhere, or Iceberg table, the shortcut still doesn’t copy the data but exposes it as OneLake Delta table. This lets you utilize Delta-specific features, such as a DirectLake semantic model without moving the data. I find this inspiring to imagine a simplified data integration in a world where one day other vendors could embrace standard file formats.</p></blockquote>
<p><strong>What about databases?</strong></p>
<p>Based on experience, a typical company has 90+ percent of its data in relational databases or connectable non-file sources, such as REST APIs and SFTP servers. In my opinion, mirroring these (sometimes huge) datasets into a file-based, pseudo-relational lakehouse rarely makes sense. Wouldn’t be nice to have shortcuts to tables in these sources and then shape and get the data you need instead of writing ETL? And even better, run cross-database queries? Wouldn’t this be a great Fabric differentiator compared to other vendors?</p>
<p>Since time immemorial, SQL Server has been supporting linked servers and heterogenous joins across databases. Then PolyBase was supposed to replace linked servers and be the Microsoft answer to broader data virtualization. Alas, both technologies didn’t make it to Fabric. Linked servers are available only in on-prem SQL Server and with limited support in Azure SQL MI. Polybase was relegated to the on-prem SQL Server.</p>
<p>I think it’s time Fabric to fulfil its zero-copy promise and say “Let me connect the dots, don’t move that data”.</p>
<p><img decoding="async" loading="lazy" src="https://prologika.com/wp-content/uploads/2017/06/060417_1725_PrologikaNe2.png" alt="" /><br />
Teo Lachev<br />
Prologika, LLC | Making Sense of Data<br />
<a href="https://prologika.com/wp-content/uploads/2016/01/logo.png" rel="attachment wp-att-12"><img decoding="async" loading="lazy" class="alignnone size-full wp-image-12" src="https://prologika.com/wp-content/uploads/2016/01/logo.png" alt="logo" width="165" height="45" /></a></p>
]]></content:encoded>
					
					<wfw:commentRss>https://prologika.com/prologika-newsletter-summer-2026/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Prologika Newsletter Spring 2026</title>
		<link>https://prologika.com/prologika-newsletter-spring-2026/</link>
					<comments>https://prologika.com/prologika-newsletter-spring-2026/#respond</comments>
		
		<dc:creator><![CDATA[Prologika - Teo Lachev]]></dc:creator>
		<pubDate>Sun, 08 Mar 2026 21:31:52 +0000</pubDate>
				<category><![CDATA[Newsletter]]></category>
		<category><![CDATA[Data Virtualization]]></category>
		<category><![CDATA[Fabric]]></category>
		<category><![CDATA[Lakehouse]]></category>
		<category><![CDATA[Power BI]]></category>
		<guid isPermaLink="false">https://prologika.com/?p=9575</guid>

					<description><![CDATA[At Ignite in November, 2025, Microsoft introduced Fabric IQ. I noted to go beyond the marketing hype and check if Fabric IQ makes any sense. The next thing I know, [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><img decoding="async" loading="lazy" class="wp-image-9535" style="padding: 0px 10px;" src="https://prologika.com/wp-content/uploads/2025/12/word-image-9533-1.png" width="120" height="120" align="left" srcset="https://prologika.com/wp-content/uploads/2025/12/word-image-9533-1.png 1024w, https://prologika.com/wp-content/uploads/2025/12/word-image-9533-1-300x300.png 300w, https://prologika.com/wp-content/uploads/2025/12/word-image-9533-1-80x80.png 80w, https://prologika.com/wp-content/uploads/2025/12/word-image-9533-1-768x768.png 768w, https://prologika.com/wp-content/uploads/2025/12/word-image-9533-1-36x36.png 36w, https://prologika.com/wp-content/uploads/2025/12/word-image-9533-1-180x180.png 180w, https://prologika.com/wp-content/uploads/2025/12/word-image-9533-1-705x705.png 705w, https://prologika.com/wp-content/uploads/2025/12/word-image-9533-1-120x120.png 120w, https://prologika.com/wp-content/uploads/2025/12/word-image-9533-1-450x450.png 450w, https://prologika.com/wp-content/uploads/2025/12/word-image-9533-1-50x50.png 50w, https://prologika.com/wp-content/uploads/2025/12/word-image-9533-1-100x100.png 100w" sizes="auto, (max-width: 120px) 100vw, 120px" /></p>
<p>At Ignite in November, 2025, Microsoft <a href="https://youtu.be/RjU0slwcZGs?si=QTRsZMg-jFQ5wsB0">introduced</a> <a href="https://youtu.be/RjU0slwcZGs?si=QTRsZMg-jFQ5wsB0">Fabric IQ</a>. I noted to go beyond the marketing hype and check if Fabric IQ makes any sense. The next thing I know, around the holidays I’m talking to an enterprise strategy manager from an airline company and McKinsey consultant about ontologies. In this newsletter I share my thoughts of FabricIQ based on my initial evaluation. Let&#8217;s start with what ontology is and how Fabric IQ uses it to integrate your enterprise data.</p>
<p><em>Ontology – A branch of philosophy, ontology is the study of being that investigates the nature of existence, the features all entities have in common, and how they are divided into basic categories of being. In computer science and AI, ontology refers to a set of concepts and categories in a subject area or domain that shows their properties and the relations between them.</em></p>
<h2>What is Fabric IQ?</h2>
<p>According to Microsoft, Fabric IQ is “a unified intelligence platform developed by Microsoft that enhances data management and decision-making through semantic understanding and AI capabilities.” Clear enough? If not, if you view Fabric as Microsoft’s answer to Palantir’s Foundry, then Fabric IQ is the Microsoft equivalent of Palantir’s Foundry Ontology, whose success apparently inspired Microsoft.</p>
<blockquote><p>Therefore, my unassuming layman definition of Fabric IQ is a metadata layer on top of data in Fabric that defines entities and their relationships so that AI can make sense of and relate the underlying data.</p></blockquote>
<p>For example, you may have an organizational semantic model built on top of an enterprise data warehouse (EDW) that spans several subject areas. And then you might have some data that isn’t in EDW and therefore outside the semantic model, such as HR file extracts in a lakehouse. You can use Fabric IQ as a glue that bridges that data together. And so, when the user asks the agent “correlate revenue by employee with hours they worked”, the agent knows where to go for answers. This screenshot shows how you can define such relationships between two entities.</p>
<p><img decoding="async" loading="lazy" src="https://learn.microsoft.com/en-us/fabric/iq/ontology/media/tutorial-1-create-ontology/semantic-model/all-entity-types.png" alt="Screenshot of the renamed entity types." /></p>
<p>Following this line of thinking, Microsoft BI practitioners may view Fabric IQ as a Power BI composite semantic model on steroids. The big difference is that a composite model can only reference other semantic models while Fabric IQ can span data in multiple formats.</p>
<h2>The Good</h2>
<p>Palantir had a head start of a decade or so compared to Microsoft Fabric, but yet even in its preview stage, I like a thing or two about Fabric IQ from what I’ve seen so far:</p>
<ul>
<li>Its oncology can span Power BI semantic models (with caveats explained in the next section), powered by best-in-class technology. As I mentioned before, this allows you to bridge all the business logic and calculations you carefully crafted in a semantic model to the rest of your Fabric data estate.</li>
<li>Fabric IQ integrates with other Microsoft technologies, such as real-time intelligence (eventhouses), Copilot Studio, Graph. This tight integration turns Fabric into a true &#8220;intelligence platform,&#8221; reducing duplicated logic, one-off models, and maintenance while enabling multi-hop reasoning and real-time operational agents.</li>
<li>Democratized and no-code friendly &#8211; Visual tools allow business users to build and evolve the ontology, lowering barriers compared to more engineering-heavy alternatives. Making it easy to use has always been a Microsoft strength.</li>
<li>Groundbreaking semantics for AI Agents: Fabric IQ elevates AI from pattern-matching to true business understanding, allowing agents to reason over cascading effects, constraints, and objectives—leading to more reliable, auditable decisions and automation.</li>
<li>Compared to Palantir, I also like that Fabric OneLake has standardized on an open Delta Parquet format and embraced data movements tools Microsoft BI pros and business users are already familiar with, such as Dataflows and pipelines, to bring data in Fabric and therefore Fabric IQ.</li>
</ul>
<h2>The Bad</h2>
<p>I hope some of these limitations will be lifted after the preview but:</p>
<ul>
<li>Only DirectLake semantic models <a href="https://learn.microsoft.com/en-us/fabric/iq/ontology/concepts-generate">are accessible</a> to AI agents. Import and DirectQuery models are not currently supported for entity and relationships binding. Not only does this limitation rule out pretty much 99.9% of the existing semantic models, but it also prevents useful business scenarios, such as accessing the data where it is with DirectQuery instead of duplicating the data in OneLake.</li>
<li>No automatic ontology building – It requires cross-functional agreement on business definitions, workshops, and governance—labor-intensive for organizations without mature semantic models. I hope Microsoft will simplify this process like how Purview has automated scans.</li>
<li>Risk of overhype vs. delivery gap – We’ve seen this before when new products got unveiled with a lot of fanfare, only to be abandoned later.</li>
</ul>
<h2>The Ugly</h2>
<p>OneLake-centric dependency. Except for shortcuts to Delta Parquet files which can be kept external, your data must be in OneLake. What about these enterprises with investments in Google BigQuery, Teradata, Snowflake, and even SQL Server or Azure SQL DB? Gotta bring that data over to OneLake. Even shortcut transformations to CSV, Parquet, JSON files in OneLake, S3, Google Cloud Storage, will copy the data to OneLake. By contrast, Palantir has limited support for virtual tables to some popular file formats, such as Parquet, Iceberg, Delta, etc.</p>
<p>What happened to all the investments in data virtualization and logical warehouses that Microsoft has made over years, such as <a href="https://prologika.com/prologika-newsletter-winter-2021/">PolyBase</a> and the deprecated <a href="https://prologika.com/synapse-serverless-the-good-the-bad-and-the-ugly/">Polaris in Synapse Serverless</a>? What’s this fascination with copying data and having all the data in OneLake? Why can’t we build Fabric IQ on top of true data virtualization?</p>
<p>Which is where I was thinking that semantic models with DirectQuery can be used as a workaround to avoid copying data over from supported data sources, but alas Fabric IQ doesn’t like them yet.</p>
<h2>Summary</h2>
<p>Microsoft Fabric IQ is a metadata layer on top of Fabric data to build ontologies and expose relevant data to AI reasoning. It will be undoubtedly appealing to enterprise customers with complex data estates and existing investments in Power BI and Fabric. However, as it stands, Fabric IQ is OneLake-centric. Expect Microsoft to invest heavily in Fabric and Fabric IQ to compete better with Palantir.</p>
<p><img decoding="async" loading="lazy" src="https://prologika.com/wp-content/uploads/2017/06/060417_1725_PrologikaNe2.png" alt="" /><br />
Teo Lachev<br />
Prologika, LLC | Making Sense of Data<br />
<a href="https://prologika.com/wp-content/uploads/2016/01/logo.png" rel="attachment wp-att-12"><img decoding="async" loading="lazy" class="alignnone size-full wp-image-12" src="https://prologika.com/wp-content/uploads/2016/01/logo.png" alt="logo" width="165" height="45" /></a></p>
]]></content:encoded>
					
					<wfw:commentRss>https://prologika.com/prologika-newsletter-spring-2026/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Prologika Newsletter Winter 2025</title>
		<link>https://prologika.com/prologika-newsletter-winter-2025/</link>
					<comments>https://prologika.com/prologika-newsletter-winter-2025/#respond</comments>
		
		<dc:creator><![CDATA[Prologika - Teo Lachev]]></dc:creator>
		<pubDate>Sat, 20 Dec 2025 00:05:32 +0000</pubDate>
				<category><![CDATA[Newsletter]]></category>
		<category><![CDATA[ETL]]></category>
		<category><![CDATA[Fabric]]></category>
		<category><![CDATA[Lakehouse]]></category>
		<guid isPermaLink="false">https://prologika.com/?p=9522</guid>

					<description><![CDATA[If Microsoft Fabric is in your future, you need to come up with a strategy to get your data in Fabric OneLake. That’s because the holy grail of Fabric is [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><img decoding="async" loading="lazy" class="wp-image-9478" style="padding: 0px 10px;" src="https://prologika.com/wp-content/uploads/2025/12/word-image-9522-1.png" alt="Diogenes holding a lantern" width="176" height="102" align="left" /></p>
<p>If Microsoft Fabric is in your future, you need to come up with a strategy to get your data in Fabric OneLake. That’s because the holy grail of Fabric is the Delta Parquet file format. The good news is that all Fabric data ingestion options (Dataflows Gen 2, pipelines, Copy Job and notebooks) support this format and the Microsoft <a href="https://learn.microsoft.com/en-us/fabric/data-warehouse/v-order">V-Order extension</a> that’s important for Direct Lake performance. Fabric also supports mirroring data from a growing list of data sources. This could be useful if your data is outside Fabric, such as EDW hosted in Google BigQuery, which is the scenario discussed in this newsletter.</p>
<h2>Avoiding mirroring issues</h2>
<p>A recent engagement required replicating some DW tables from Google BigQuery to a Fabric Lakehouse. We considered the <a href="https://learn.microsoft.com/en-us/fabric/mirroring/google-bigquery">Fabric mirroring feature for Google BigQuery</a> (back then in private preview, now in public preview) and learned some lessons along the way:</p>
<p>1. 400 Error during replication configuration – Caused by attempting to use a read-only GBQ dataset that is linked to another GBQ dataset, but the link was broken.</p>
<p>2. Internal System Error – Again caused by GBQ linked datasets which are read-only. Fabric mirroring requires GBQ change history to be enabled on tables so that it can track changes and only mirror incremental changes after first initial load.</p>
<p>3. (Showstopper for this project) The two permissions that raised security red flags are bigquery.datasets.create and bigquery.jobs.create. To grant those permissions, you must assign one of these BigQuery roles:</p>
<p>• BigQuery Admin</p>
<p>• BigQuery Data Editor</p>
<p>• BigQuery Data Owner</p>
<p>• BigQuery Studio Admin</p>
<p>• BigQuery User</p>
<p>All these roles grant other permissions, and the client was cautious about data security. At the end, we end up using a nightly Fabric Copy Job to replicate the data.</p>
<h2>Fabric Copy Job Pros and Cons</h2>
<p>The client was overall pleased with the Fabric Copy Job.</p>
<p><strong>Pros</strong></p>
<ul>
<li>250 million rows replicated in 30-40 seconds!</li>
<li>You can have only one job to replicate all tables in Overwrite mode.</li>
<li>In the simplest case, you don’t need to create pipelines.</li>
</ul>
<p><strong>Cons </strong></p>
<p>The Copy Job is work in progress and subject to various limitations.</p>
<ul>
<li>No incremental extraction</li>
<li>You can’t mix different load options (Append and Overwrite) so you must split tables in separate jobs</li>
<li>No custom SQL SELECT when copying multiple tables</li>
<li>(Bug) Lost explicit column bindings when making changes</li>
<li>Cannot change the job’s JSON file</li>
<li>The user interface is clunky and it’s difficult to work with</li>
<li>No failure notification mechanism. As a workaround: add Copy Job to data pipeline or call it via REST API</li>
</ul>
<h2>Summary</h2>
<p>In summary, the Fabric Google BigQuery built-in mirroring could be useful for real-time data replication. However, it relies on GBQ change history which requires certain permissions. Kudos to Microsoft for their excellent support during the private preview.</p>
<p><img decoding="async" loading="lazy" src="https://prologika.com/wp-content/uploads/2017/06/060417_1725_PrologikaNe2.png" alt="" /><br />
Teo Lachev<br />
Prologika, LLC | Making Sense of Data<br />
<a href="https://prologika.com/wp-content/uploads/2016/01/logo.png" rel="attachment wp-att-12"><img decoding="async" loading="lazy" class="alignnone size-full wp-image-12" src="https://prologika.com/wp-content/uploads/2016/01/logo.png" alt="logo" width="165" height="45" /></a></p>
]]></content:encoded>
					
					<wfw:commentRss>https://prologika.com/prologika-newsletter-winter-2025/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Prologika Newsletter Fall 2025</title>
		<link>https://prologika.com/prologika-newsletter-fall-2025/</link>
					<comments>https://prologika.com/prologika-newsletter-fall-2025/#respond</comments>
		
		<dc:creator><![CDATA[Prologika - Teo Lachev]]></dc:creator>
		<pubDate>Mon, 22 Sep 2025 16:48:33 +0000</pubDate>
				<category><![CDATA[Newsletter]]></category>
		<category><![CDATA[Fabric]]></category>
		<category><![CDATA[Semantic Model]]></category>
		<guid isPermaLink="false">https://prologika.com/?p=9476</guid>

					<description><![CDATA[Like the Ancient Greek philosopher Diogenes, who walked the streets of Athens with a lamp to find one honest man, I have been searching for a convincing Fabric feature for [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><img decoding="async" loading="lazy" class="wp-image-9478" style="padding: 0px 10px;" src="https://prologika.com/wp-content/uploads/2025/09/a-person-holding-a-lantern-ai-generated-content-m.png" alt="Diogenes holding a lantern" width="95" height="127" align="left" srcset="https://prologika.com/wp-content/uploads/2025/09/a-person-holding-a-lantern-ai-generated-content-m.png 720w, https://prologika.com/wp-content/uploads/2025/09/a-person-holding-a-lantern-ai-generated-content-m-225x300.png 225w, https://prologika.com/wp-content/uploads/2025/09/a-person-holding-a-lantern-ai-generated-content-m-529x705.png 529w, https://prologika.com/wp-content/uploads/2025/09/a-person-holding-a-lantern-ai-generated-content-m-450x600.png 450w" sizes="auto, (max-width: 95px) 100vw, 95px" /> Like the Ancient Greek philosopher Diogenes, who walked the streets of Athens with a lamp to find one honest man, I have been searching for a convincing Fabric feature for my clients. As Microsoft Fabric evolves, more scenarios unfold. For example, Direct Lake storage mode could help you alleviate memory pressure with large semantic models in certain scenarios, as it did for one client. This newsletter summarizes the important takeaways from this project. If this sounds interesting and you are geographically close to Atlanta, I invite you to the December 1st meeting of the <a href="https://www.meetup.com/atlanta-microsoft-business-intelligence-users">Atlanta MS BI Group</a> where I&#8217;ll present the implementation details.</p>
<h1>About the project</h1>
<p>In this case, the client had a 40 GB semantic model with 250 million rows spread across two fact tables. The semantic model imported data from a Google BigQuery (GBQ) data warehouse. The client applied every trick in the book to optimize the model, but they’ve found themselves forced to upgrade from a Power BI F64 to F128 to F256 capacity.</p>
<blockquote><p>I’ve written in the past about my frustration with Power BI/Fabric capacity resource limits. While the 25 GB RAM grant of a P1/F64 capacity for each dataset is generous for smaller semantic models, such as for self-service BI, it’s inadequate for large organizational semantic models. Ultimately, the developer must face gut wrenching decisions, such as whether to split the model into smaller semantic models to obey what are in my opinion artificially low and inflexible memory limits or ask for more money.</p></blockquote>
<p>We’ve decided to replicate the GBQ data to a Fabric lakehouse and try Direct Lake to avoid the dataset refresh, which requires at least twice the memory. Granted, replicating data is an awkward solution, but currently Direct Lake requires data to be in Delta tables (Fabric Lakehouse, Data Warehouse, or shortcuts to Delta tables, such as in OneLake or Databricks).</p>
<p>Next, we migrated the largest semantic model from import to Direct Lake. You can find the technical details for the replication and migration steps we took in my blog “<a href="https://prologika.com/migrating-fabric-import-semantic-models-to-direct-lake/">Migrating Fabric Import Semantic Models to Direct Lake</a>”.</p>
<h1>Performance considerations</h1>
<p>The following screenshot is taken from the Fabric Capacity Metrics app and it shows the maximum metrics over 14 days. The two enclosed items of interest are the original imported semantic model (the first item on the list) and its DL counterpart (the seventh item on the list).</p>
<p>The Direct Lake memory utilization was at a par with the imported model. With 1/5 of the user audience testing the dataset in production environment, that dataset grew to a maximum of 25 GB memory utilization which is in line with the imported model. It could have been interesting to downgrade the capacity, such as to F64, and observe how the DL model would react to memory pressure. However, as shown in the screenshot, the client had other large semantic models that can exhaust the F64 25 memory grant so we couldn’t perform this test.</p>
<p><img decoding="async" width="1277" height="294" loading="lazy" class="wp-image-9479" src="https://prologika.com/wp-content/uploads/2025/09/a-screenshot-of-a-computer-ai-generated-content-ma.png" alt="A screenshot of a computer AI-generated content may be incorrect." srcset="https://prologika.com/wp-content/uploads/2025/09/a-screenshot-of-a-computer-ai-generated-content-ma.png 1277w, https://prologika.com/wp-content/uploads/2025/09/a-screenshot-of-a-computer-ai-generated-content-ma-300x69.png 300w, https://prologika.com/wp-content/uploads/2025/09/a-screenshot-of-a-computer-ai-generated-content-ma-1030x237.png 1030w, https://prologika.com/wp-content/uploads/2025/09/a-screenshot-of-a-computer-ai-generated-content-ma-768x177.png 768w, https://prologika.com/wp-content/uploads/2025/09/a-screenshot-of-a-computer-ai-generated-content-ma-705x162.png 705w, https://prologika.com/wp-content/uploads/2025/09/a-screenshot-of-a-computer-ai-generated-content-ma-450x104.png 450w" sizes="auto, (max-width: 1277px) 100vw, 1277px" /></p>
<blockquote><p>Again, what we are saving here is the additional memory required for refreshing the model. In a sense, we shifted the model refresh to replicating the data from Google Big Query to a Fabric lakehouse. On the downside, an error during the replication process could leave the replicated tables in an inconsistent state (and user complaints because reports would show no data or stale data) whereas a failure during refreshing the model would fall back on the old model (Fabric builds a new in-memory cache during model refreshing).</p></blockquote>
<p>We didn’t witness excessive CPU pressure during production testing. Further, the team didn’t notice any report performance degradation or increased CU capacity utilization.</p>
<h1>Summary</h1>
<p>Assuming you have exhausted traditional methods to alleviate memory pressure, such eliminating high-cardinality column, incremental refresh, etc., Direct Lake is a viable option to conserve memory of Fabric semantic models. However, it may require replicating your data to a Fabric lakehouse or migrating your data warehouse to Fabric so that it uses Fabric storage (Delta Parquet format) required for Direct Lake. If this is a new project and you expect large semantic models, your architecture should strongly consider Fabric Data Warehouse or Lakehouse to take advantage of Direct Lake storage.</p>
<p><img decoding="async" loading="lazy" src="https://prologika.com/wp-content/uploads/2017/06/060417_1725_PrologikaNe2.png" alt="" /><br />
Teo Lachev<br />
Prologika, LLC | Making Sense of Data<br />
<a href="https://prologika.com/wp-content/uploads/2016/01/logo.png" rel="attachment wp-att-12"><img decoding="async" loading="lazy" class="alignnone size-full wp-image-12" src="https://prologika.com/wp-content/uploads/2016/01/logo.png" alt="logo" width="165" height="45" /></a></p>
]]></content:encoded>
					
					<wfw:commentRss>https://prologika.com/prologika-newsletter-fall-2025/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Prologika Newsletter Summer 2025</title>
		<link>https://prologika.com/prologika-newsletter-summer-2025/</link>
					<comments>https://prologika.com/prologika-newsletter-summer-2025/#respond</comments>
		
		<dc:creator><![CDATA[Prologika - Teo Lachev]]></dc:creator>
		<pubDate>Mon, 16 Jun 2025 19:28:45 +0000</pubDate>
				<category><![CDATA[Newsletter]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[LLM]]></category>
		<guid isPermaLink="false">https://prologika.com/?p=9423</guid>

					<description><![CDATA[The May release of Power BI Desktop added a feature called Translytical Task Flows which aims to augment Power BI reports with rudimentary writeback capabilities, such as to make corrections [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><img decoding="async" loading="lazy" class="wp-image-9297 " style="padding: 0px 10px;" src="https://prologika.com/wp-content/uploads/2025/05/thumbnail-image-a-digital-illustration-of-a-comput-1.jpeg" alt="computer memory with queries executing and Microsoft Fabric logo. Image 4 of 4" width="137" height="120" align="left" /></p>
<p>The May release of Power BI Desktop added a feature called <a href="https://powerbi.microsoft.com/en-us/blog/power-bi-may-2025-feature-summary/#post-29934-_Toc485152301">Translytical Task Flows</a> which aims to augment Power BI reports with rudimentary writeback capabilities, such as to make corrections to data behind a report. Previously, one way to accomplish this was to integrate the report with Power Apps as I demonstrated a while back <a href="https://prologika.com/power-bi-writeback/">here</a>. My claim to fame was that Microsoft liked this demo so much that it was running for years on big monitors in the local Microsoft office!</p>
<p>Are translytical flows a better way to implement report writeback? I followed the <a href="https://learn.microsoft.com/en-us/power-bi/create-reports/translytical-task-flow-tutorial">steps</a> to test this feature and here are my thoughts.</p>
<h1>The Good</h1>
<p>I like that translytical flows don’t require external integration and additional licensing. By contrast, the Power Apps integration required implementing an app and incurring additional licensing cost.</p>
<p>I like that Microsoft is getting serious about report writeback and has extended Power BI Desktop with specific features to support it, such as action buttons, new slicers, and button-triggered report refresh.</p>
<p>I like that you can configure the action button to refresh the report after writeback so you can see immediately the changes (assuming DirectQuery or DirectLake semantic models). I tested the feature with a report connected to a published dataset and it works. Of course, if the model imports data, refreshing the report won’t show the latest.</p>
<h1>The Bad</h1>
<p>Currently, you must use the new button, text, or list slicers, which can only provide a single value to the writeback function. No validation, except if you use a list slicer. From end user experience, every modifiable field would need a separate textbox. This is horrible UI! Ideally, I’d like to see Microsoft extending the Table visual to allow editing in place.</p>
<p>Translytical flows require a Python function to propagate the changes. Although this opens new possibilities (you can do more things with custom code), what happened to low-code, no-code mantra?</p>
<h1>The Ugly</h1>
<p>Currently, the data can be written to only four destinations:</p>
<ul>
<li>Fabric SQL DB (Azure SQL DB provisioned in Fabric)</li>
<li>Fabric Lakehouse</li>
<li>Fabric Warehouse</li>
<li>Fabric Mirrored DB</li>
</ul>
<p>Notice the “Fabric” prefix in all four options? Customers will probably interpret this as another attempt to force them into Fabric. Not to mention that this imitation excludes 99% of real-life scenarios where the data is in other data sources. I surely hope Microsoft will open this feature to external data sources in future.</p>
<p>So, is the Power Apps integration for writeback obsolete? Not really because it is more flexible and provides better user experience at the expense of additional licensing cost.</p>
<blockquote><p>In summary, as they stand today, transalytical task flows attempt to address basic writeback needs within Fabric, such as changing a limited number of report fields or performing massive updates on a single field. They are heavily dependent on Fabric and support writeback to only Fabric data sources. Keep an eye on this feature with the hope that it will evolve over time to something more useful.</p></blockquote>
<p>&nbsp;</p>
<p><img decoding="async" loading="lazy" src="https://prologika.com/wp-content/uploads/2017/06/060417_1725_PrologikaNe2.png" alt="" /><br />
Teo Lachev<br />
Prologika, LLC | Making Sense of Data<br />
<a href="https://prologika.com/wp-content/uploads/2016/01/logo.png" rel="attachment wp-att-12"><img decoding="async" loading="lazy" class="alignnone size-full wp-image-12" src="https://prologika.com/wp-content/uploads/2016/01/logo.png" alt="logo" width="165" height="45" /></a></p>
]]></content:encoded>
					
					<wfw:commentRss>https://prologika.com/prologika-newsletter-summer-2025/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Prologika Newsletter Spring 2025</title>
		<link>https://prologika.com/prologika-newsletter-spring-2025/</link>
					<comments>https://prologika.com/prologika-newsletter-spring-2025/#respond</comments>
		
		<dc:creator><![CDATA[Prologika - Teo Lachev]]></dc:creator>
		<pubDate>Sat, 15 Mar 2025 15:36:09 +0000</pubDate>
				<category><![CDATA[Newsletter]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[LLM]]></category>
		<guid isPermaLink="false">https://prologika.com/?p=9372</guid>

					<description><![CDATA[Looking for easy ways to create intelligent bots or Retrieval-Augmented Generation (RAG) apps? Microsoft Copilot Studio should help. After a year of letting it (and other Microsoft LLM offerings) simmer [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><img decoding="async" loading="lazy" class="wp-image-9297 " style="padding: 0px 10px;" src="https://prologika.com/wp-content/uploads/2025/02/AI-intelligent-bot.jpg" alt="computer memory with queries executing and Microsoft Fabric logo. Image 4 of 4" width="137" height="120" align="left" />Looking for easy ways to create intelligent bots or Retrieval-Augmented Generation (RAG) apps? <a href="https://www.microsoft.com/en-us/microsoft-copilot/microsoft-copilot-studio?msockid=10fbdb3804c7683f32a7ce7c05116900">Microsoft Copilot Studio</a> should help. After a year of letting it (and other Microsoft LLM offerings) simmer and inspired by the latest hoopla from the Ignite conference, I took another look at Microsoft Copilot Studio and share my findings in this newsletter. For the uninitiated, Copilot Studio lets you implement AI-powered smart bots (“agents”) for deriving knowledge from documents or websites. Basically, you can view the relationship of Copilot Studio to <a href="https://en.wikipedia.org/wiki/Retrieval-augmented_generation">Retrieval-augmented Generation</a> apps as what Power BI is to self-service BI.</p>
<p>&nbsp;</p>
<p>Copilot Studio licensing starts at $200 per month for up to 25,000 messages (interactions between user and agent) although at Ignite Microsoft hinted that pay-as-you-go licensing will be coming.</p>
<p><strong>The Good</strong></p>
<p>A few months ago, when I <a href="https://prologika.com/llm-adventures-rag-apps/">discussed</a> RAG apps, it was obvious that a lot of custom code had to be written to glue the services together and implement the user interface. Microsoft Copilot Studio has the potential to change and simplify this. It offers a Power Automate-like environment for no-code, low-code implementation of AI agents and therefore opens new possibilities for faster implementation of various and specialized AI agents across the enterprise. I was impressed by how easy the process was and how capable the tool was to create more complex topics, such as conditional branches based on user input.</p>
<p>Like Power BI, the tool gets additional appeal from its integration with the Microsoft ecosystem. For example, it can index SharePoint and OneDrive documents. It can integrate with Power Automate, Azure AI Search and Azure Open AI.</p>
<blockquote><p>I was impressed by how easy is to use the tool to connect to and intelligently search an existing website. For now, I see this as being its main strength. Organizations can quickly implement agents to help their employees or external users to derive knowledge from intranet or Internet websites.</p></blockquote>
<p>To demonstrate this, I implemented an agent to index my blog and embedded it below for you to try it out before my free trial expires. Please feel free to ask more sophisticated questions, such as “What’s the author’s sentiment toward Fabric?”, &#8220;What are the pros and cons of Fabric?&#8221;. Or “I need help with Power BI budget” (I got innovative here and implemented a conditional topic with branches depending on the budget you specify). I instructed the tool to stay only within the content of  my website, so the answers are not diluted from other public sources. Given that no custom code was written, Copilot Studio is pretty impressive.</p>
<p><iframe style="width: 400px; height: 600px;" src="https://copilotstudio.microsoft.com/environments/Default-e7b81d0a-a949-4103-83dc-feff6277c109/bots/cr534_prologikaWebsiteQACopilot/webchat?__version__=2" frameborder="0"><span data-mce-type="bookmark" style="display: inline-block; width: 0px; overflow: hidden; line-height: 0;" class="mce_SELRES_start">﻿</span><span data-mce-type="bookmark" style="display: inline-block; width: 0px; overflow: hidden; line-height: 0;" class="mce_SELRES_start">﻿</span></iframe></p>
<p><strong>The Bad</strong></p>
<p>Everyone wants to be autonomous and AI agents are no exception. In fact, &#8220;<strong>autonomous agent&#8221;</strong> is the buzzword of AI world today. Not to be outdone, Copilot Studio claims that it can “build agents that operate independently to dynamically plan, learn, and escalate on your behalf”. However, as the tool stands today, I don’t think there is much to this claim. Or it could be that my definition of “autonomous” is different than Microsoft’s.</p>
<p>To me, an autonomous agent must be capable of making decisions and taking actions on its own. Like you tell your assistant that you plan a trip, give her some constraints, such as how much to spend on hotel and air, and let her make travel reservations. As it stands, Copilot Studio offers none of this. It follows a workflow you specify. Again, its output is more or less a smarter bot than the ones you see on many websites.</p>
<p>However, at Ignite Microsoft claimed that autonomy is coming so it will be interesting to see how the tool will evolve. Don’t get me wrong. Even as it stands, I believe the tool has enormous potential for more intelligent search and retrieval of information.</p>
<p><strong>The Ugly</strong></p>
<p>My basic complaint as of now is performance. It took the tool 10 minutes to index a PDF document. Then in a momentary lapse of reason, I connected it to an Azure SQL Database with Adventure Works with 15 tables (the max number of tables currently supported) and it’s still not done indexing after a day. Given that many AI implementations would require searching the data in relational databases, this is not acceptable. Not to mention there isn&#8217;t much insight on how far it&#8217;s done indexing or limit the number of fields it should index.</p>
<blockquote><p>Therefore, I believe most real-world architectures for implementing AI agents will take the path Copilot Studio-&gt;Azure AI Search -&gt;Azure Open AI, where Copilot Studio is used for implementing the UI and workflows (topics and actions), while the data indexing is done by Azure AI Search with semantic ranking in conjunction with Azure Open AI for embedded vectors.</p></blockquote>
<p>&nbsp;</p>
<p><img decoding="async" loading="lazy" src="https://prologika.com/wp-content/uploads/2017/06/060417_1725_PrologikaNe2.png" alt="" /><br />
Teo Lachev<br />
Prologika, LLC | Making Sense of Data<br />
<a href="https://prologika.com/wp-content/uploads/2016/01/logo.png" rel="attachment wp-att-12"><img decoding="async" loading="lazy" class="alignnone size-full wp-image-12" src="https://prologika.com/wp-content/uploads/2016/01/logo.png" alt="logo" width="165" height="45" /></a></p>
]]></content:encoded>
					
					<wfw:commentRss>https://prologika.com/prologika-newsletter-spring-2025/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Prologika Newsletter Winter 2024</title>
		<link>https://prologika.com/prologika-newsletter-winter-2024/</link>
					<comments>https://prologika.com/prologika-newsletter-winter-2024/#respond</comments>
		
		<dc:creator><![CDATA[Prologika - Teo Lachev]]></dc:creator>
		<pubDate>Sat, 14 Dec 2024 15:53:28 +0000</pubDate>
				<category><![CDATA[Newsletter]]></category>
		<category><![CDATA[Fabric]]></category>
		<category><![CDATA[Power BI]]></category>
		<guid isPermaLink="false">https://prologika.com/?p=9337</guid>

					<description><![CDATA[I conducted recently an assessment for a client facing memory pressure in Power BI Premium. You know these pesky out of memory errors when refreshing a biggish dataset. They started [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><img decoding="async" loading="lazy" class="wp-image-9297 " style="padding: 0px 10px;" src="https://prologika.com/wp-content/uploads/2024/08/computer-memory-with-queries-executing-and-microso.jpeg" alt="computer memory with queries executing and Microsoft Fabric logo. Image 4 of 4" width="137" height="120" align="left" />I conducted recently an assessment for a client facing memory pressure in Power BI Premium. You know these pesky out of memory errors when refreshing a biggish dataset. They started with P1, moved to P2, and now are on P3 but still more memory is needed to satisfy the memory appetite of full refresh. The runtime memory footprint of the problematic semantic model with imported data is 45 GB and they’ve done their best to optimize it. This newsletter outlines a few strategies to tackle excessive memory consumption with large semantic models. Unfortunately, given the current state of Power BI boxed capacities, no option is perfect and at end a compromise will probably be needed somewhere between latency and performance.</p>
<h4><strong>Why I don’t like Premium licensing</strong></h4>
<p>Since its beginning, Power BI Pro per-user licensing (and later Premium Per User (PPU) licensing) has been very attractive. Many organizations with a limited number of report users flocked to Power BI to save cost. However, organizations with more BI consumers gravitated toward premium licensing where they could have unlimited number of report readers against a fixed monthly fee starting at listed price of $5,000/mo for P1. Sounds like a great deal, right?</p>
<p>I must admit that I detest the premium licensing model because it boxes into certain resource constraints, such as 8 backend cores and 25 GB RAM for P1. There are no custom configurations to let you balance between compute and memory needs. And while there is an auto-scale compute model, it’s very coarse and it applies only to processing cores. The memory constraints are especially problematic given that imported models are memory resident and require more than twice the memory for full refresh. From the outside, these memory constraints seem artificially low to force clients into perpetual upgrades. The new Fabric F capacities that supersede the P plans are even more expensive, justifying the price increase with the added flexibility to pause the capacity which is often impractical.</p>
<p>It looks to me that the premium licensing is pretty good deal for Microsoft. Outgrown 25 GB of RAM in P1? Time to shelve another 5K per month for 25 GB more even if you don’t need more compute power. Meanwhile, the price of 32GB of RAM is less than $100 and falling.</p>
<blockquote><p>It will be great if at some point Power BI introduces custom capacities. Even better, how about auto-scaling where the capacity resources (both memory and CPU) scale up and down on demand within minutes, such as adding more memory during refresh and reducing the memory when the refresh is over?</p></blockquote>
<p><strong>Strategies to combat out-of-memory scenarios</strong></p>
<p>So, what should you do if you are strapped for cash? Consider evaluating and adopting one or more of the following memory saving techniques, including:</p>
<ul>
<li>Switching to PPU licensing with a limited number of report users. PPU is equivalent of P3 and grants 100GB RAM per dataset.</li>
<li>Optimizing aggressively the model storage when possible, such as removing high-cardinality columns</li>
<li>Configuring aggressive incremental refresh policies with polling expressions</li>
<li>Moving large fact tables to a separate semantic model (remember that the memory constraints are per dataset and not across all the datasets in the capacity)</li>
<li>Implementing DirectQuery features, such as composite models and <a href="https://prologika.com/power-bi-hybrid-tables/"><strong>hybrid tables</strong></a></li>
<li>Switching to a hybrid architecture with on-prem semantic model(s) hosted in SQL Server Analysis Services where you can control the hardware configuration and you’re not charge for more memory.</li>
<li>Lobbying Microsoft for much larger memory limits or to bring your own memory (good luck with that but it might be an option if you work for a large and important company)</li>
</ul>
<p><strong>Considering Direct Lake storage</strong></p>
<p>If Fabric is in your future, one relatively new option to tackle out-of-memory scenarios that deserves to be evaluated and added to the list is semantic models configured for <a href="https://learn.microsoft.com/en-us/fabric/get-started/direct-lake-overview">Direct Lake</a> storage. Direct Lake on-demand loading should utilize memory much more efficiently for interactive operations, such as Power BI report execution. This is a bonus to the fact that data Direct Lake models don’t require refresh. Eliminating refresh could save tremendous amount of memory to start with, even if you apply advanced techniques such as incremental refresh or hybrid tables to models with imported data.</p>
<p>I did limited testing to compare performance of import and Direct Lake and posted detailed results in the “<a href="https://prologika.com/fabric-direct-lake-memory-utilization/">Fabric Direct Lake: Memory Utilization with Interactive Operations</a>” blog.</p>
<blockquote><p>I concluded that if Direct Lake is an option for you, it should be at the forefront of your efforts to combat out-of-memory errors with large datasets.</p></blockquote>
<p>On the downside, more than likely you’ll have to implement ETL processes to synchronize your data warehouse to a Fabric lakehouse, unless your data is in Fabric to start with, or you use Fabric database mirroring for the currently supported data sources (Azure SQL DB, Cosmos, and Snowflake). I’m not counting the data synchronization time as a downside.</p>
<p><img decoding="async" loading="lazy" src="https://prologika.com/wp-content/uploads/2017/06/060417_1725_PrologikaNe2.png" alt="" /><br />
Teo Lachev<br />
Prologika, LLC | Making Sense of Data<br />
<a href="https://prologika.com/wp-content/uploads/2016/01/logo.png" rel="attachment wp-att-12"><img decoding="async" loading="lazy" class="alignnone size-full wp-image-12" src="https://prologika.com/wp-content/uploads/2016/01/logo.png" alt="logo" width="165" height="45" /></a></p>
]]></content:encoded>
					
					<wfw:commentRss>https://prologika.com/prologika-newsletter-winter-2024/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Prologika Newsletter Fall 2024</title>
		<link>https://prologika.com/prologika-newsletter-fall-2024/</link>
					<comments>https://prologika.com/prologika-newsletter-fall-2024/#respond</comments>
		
		<dc:creator><![CDATA[Prologika - Teo Lachev]]></dc:creator>
		<pubDate>Fri, 13 Sep 2024 18:06:32 +0000</pubDate>
				<category><![CDATA[Newsletter]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[DAX]]></category>
		<category><![CDATA[SQL Server]]></category>
		<guid isPermaLink="false">https://prologika.com/?p=9296</guid>

					<description><![CDATA[When it comes to Generative AI and Large Language Models (LLMs), most people fall into two categories. The first is alarmists. These people are concerned about the negative connotations of [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><img decoding="async" loading="lazy" class="wp-image-9297" style="padding: 0px 10px;" src="https://prologika.com/wp-content/uploads/2024/09/text-to-sql-copilot-image-3-of-4.jpeg" alt="Text to SQL copilot" width="137" height="116" align="left" /> When it comes to Generative AI and Large Language Models (LLMs), most people fall into two categories. The first is alarmists. These people are concerned about the negative connotations of indiscriminate usage of AI, such as losing their jobs or military weapons for mass annihilation. The second category are deniers, and I must admit I was one of them. When Generative AI came out, I dismissed it as vendor propaganda, like Big Data, auto-generative BI tools, lakehouses, ML, and the like. But the more I learn and use Generative AI, the more credit I believe it deserves. Because LLMs are trained with human and programming languages, one natural case where they could be helpful are code copilots, which is the focus of this newsletter. Let&#8217;s give Generative AI some credit!</p>
<h1><strong>Text2SQL</strong></h1>
<p>I have to say that I was impressed with LLM. I used the excellent <a href="https://github.com/RicZhou-MS/Text2SQL"><strong>Ric Zhou’s Text2SQL sample</strong></a> as a starting point inside Visual Studio Code.</p>
<p><img decoding="async" loading="lazy" class="alignnone size-full wp-image-9256" src="https://prologika.com/wp-content/uploads/2024/07/a-screenshot-of-a-chat-description-automatically-2.png" alt="A screenshot of a chat Description automatically generated" width="1430" height="735" srcset="https://prologika.com/wp-content/uploads/2024/07/a-screenshot-of-a-chat-description-automatically-2.png 1430w, https://prologika.com/wp-content/uploads/2024/07/a-screenshot-of-a-chat-description-automatically-2-300x154.png 300w, https://prologika.com/wp-content/uploads/2024/07/a-screenshot-of-a-chat-description-automatically-2-1030x529.png 1030w, https://prologika.com/wp-content/uploads/2024/07/a-screenshot-of-a-chat-description-automatically-2-768x395.png 768w, https://prologika.com/wp-content/uploads/2024/07/a-screenshot-of-a-chat-description-automatically-2-705x362.png 705w, https://prologika.com/wp-content/uploads/2024/07/a-screenshot-of-a-chat-description-automatically-2-450x231.png 450w" sizes="auto, (max-width: 1430px) 100vw, 1430px" /></p>
<p>The sample uses the Python streamlit framework to create a web app that submits natural questions to Azure OpenAI. I was amazed how simple the LLM input was. Given that it’s trained with many popular languages, including SQL, all you have to do is provide some context, database schema (generated in a simple format by a provided tool), and a few prompts:</p>
<pre>[{'role': 'system', 'content': 'You are smart SQL expert who can do text to SQL, following is the Azure SQL database data model &lt;database schema&gt;},
{'role': 'user', 'content': 'What are the three best selling cities for the "AWC Logo Cap" product?'},
{'role': 'assistant', 'content': 'SELECT TOP 3 A.City, sum(SOD.LineTotal) AS TotalSales \\nFROM [SalesLT].[SalesOrderDe...BY A.City \\nORDER BY TotalSales DESC; \\n'},
{'role': 'user', 'content': natural question here'}]</pre>
<p>Let’s drill into these prompts.</p>
<ol>
<li>The first prompt is an example of role prompting which provides context to the model for our intention to act as a SQL expert.</li>
<li>The sample includes a tool that generates the database schema consisting of table and columns in the following format below. Notice that referential integrity constraints are not included (the model doesn’t need to know how the tables are related!)</li>
<li>The next “few shot” prompt assumes the role of an end user who will ask a natural question, such as ‘What are the three best selling cities for the “AWC Logo Cap” product? This is followed by the assistant’s response who hints the model what the correct query should be.</li>
<li>Then the request for new query follows.</li>
</ol>
<p>You can find more Text2SQL implementation details in my post “<a href="https://prologika.com/llm-adventures-text2sql/">LLM Adventures: Text2SQL</a>”.</p>
<h1><strong>Text2DAX</strong></h1>
<p>As a Microsoft BI practitioner, the next natural stop was Text2DAX. But wait, we have a Microsoft Fabric Copilot already for this, right? Yes, but what happens when you click the magic button in PBI Desktop? You are greeted that you need to purchase F64 or larger capacity. It’s a shame that Microsoft has decided that AI should be a super-premium feature. Given this horrible predicament, what would an innovative developer strapped for cash do? Create their own copilot of course!</p>
<p>Building upon the previous sample, this is remarkably simple. First, I obtained the model schema using Analysis Services Data Management Views (DMVs). Second, I changed slightly the prompt, along these lines:</p>
<p><em>“As a DAX expert, examine a Power BI data model with the following table definitions where TableName lists the tables and ColumnName lists the columns, and provide DAX query that can answer user questions based on the data model, you will think step by step throughout and return DAX statement directly without any additional explanation.”</em></p>
<p>You can find more Text2DAX implementation details in my post “<a href="https://prologika.com/llm-adventures-text2dax/">LLM Adventures: Text2DAX</a>”.</p>
<blockquote><p>In summary, it appears that LLM can effectively assist us in writing code. The emphasis is on <strong>assist</strong> because I view the LLM role as a second set of eyes. Hey, what do you think about this problem I’m trying to solve here? LLM doesn’t absolve us from doing our homework and learning the fundamentals, nor it can compensate for improper design. While LLM might not always generate the optimum code and might sometimes fabricate, it can definitely assist you in creating business calculations, generating test queries, and learning along the way.</p></blockquote>
<p>BTW, you can use any of the publicly available LLM apps, such as Copilot, ChatGPT, Google Gemini or Perplexity (you don’t need the sample app I’ve demonstrated) for Text2SQL and Text2DAX and probably you will obtain similar results if you give it the right prompts. I took this approach because I was interested in automating the process for business users.</p>
<p><img decoding="async" loading="lazy" src="https://prologika.com/wp-content/uploads/2017/06/060417_1725_PrologikaNe2.png" alt="" /><br />
Teo Lachev<br />
Prologika, LLC | Making Sense of Data<br />
<a href="https://prologika.com/wp-content/uploads/2016/01/logo.png" rel="attachment wp-att-12"><img decoding="async" loading="lazy" class="alignnone size-full wp-image-12" src="https://prologika.com/wp-content/uploads/2016/01/logo.png" alt="logo" width="165" height="45" /></a></p>
]]></content:encoded>
					
					<wfw:commentRss>https://prologika.com/prologika-newsletter-fall-2024/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Prologika Newsletter Summer 2024</title>
		<link>https://prologika.com/prologika-newsletter-summer-2024/</link>
					<comments>https://prologika.com/prologika-newsletter-summer-2024/#respond</comments>
		
		<dc:creator><![CDATA[Prologika - Teo Lachev]]></dc:creator>
		<pubDate>Sun, 16 Jun 2024 15:11:21 +0000</pubDate>
				<category><![CDATA[Newsletter]]></category>
		<category><![CDATA[Data Warehousing]]></category>
		<category><![CDATA[Fabric]]></category>
		<category><![CDATA[Lakehouse]]></category>
		<guid isPermaLink="false">https://prologika.com/?p=9234</guid>

					<description><![CDATA[I’ve written in the past about the dangers of blindly following “modern” data architectures (see the “Are you modern yet?” and “Data Lakehouse: The Good, the Bad, and the Ugly”) [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><img decoding="async" loading="lazy" class="" style="padding: 0px 10px;" src="https://prologika.com/wp-content/uploads/2024/06/lake.png" alt="" width="112" height="99" align="left" />I’ve written in the past about the dangers of blindly following “modern” data architectures (see the “<a href="https://prologika.com/are-you-modern-yet/">Are you modern yet?</a>” and “<a href="https://prologika.com/data-lakehouse-the-good-the-bad-and-the-ugly/">Data Lakehouse: The Good, the Bad, and the Ugly</a>”) but a recent assessment inspired to me write about this topic again. This newsletter advocates a hybrid and cautionary approach for data integration to avoid overdoing data lakes and warns about pitfalls of over-staging source data to files. It recommends instead following the &#8220;Discipline at the core, flexibility at the edge&#8221; methodology with emphasis on implementing enterprise data warehouse and organizational semantic models.</p>
<h1>Data Lake Overstaging</h1>
<p>How did the large vendor attempt to solve these horrible issues? Modern Data Warehouse (MDM) architecture of course. Nothing wrong with it except that EDW and organizational semantic model(s) are missing and that most of the effort went into implementing the data lake medallion architecture where all the incoming data ended up staged as Parquet files. It didn’t matter that 99% of the data came from relational databases. Further, to solve a data change tracking requirement, the vendor decided to create a new file each time ETL runs. So even if nothing has changed in the source feed, the data is duplicated should one day the user wants to go back in time and see what the data looked like then. There are of course better ways to handle this that doesn’t even require ETL, such as SQL Server temporal tables, but I digress.</p>
<p>At least some cool heads prevailed and the Silver layer got implemented as a relational ODS to serve the needs of home-grown applications, so the apps didn’t have to deal with files. What about EDW and organizational semantic models? Not there because the project ran out of budget and time. I bet if that vendor got hired today, they would have gone straight for Fabric Lakehouse and Fabric premium pricing (nowadays Microsoft treats partners as an extension to its salesforce and requires them to meet certain revenue targets as I explain in “<a href="https://prologika.com/dissolving-partnerships/">Dissolving Partnerships</a>“), which alone would have produced the same outcome.</p>
<blockquote><p>What did the vendor accomplish? Not much. Nor only didn’t the implementation address the main challenges, but it introduced new, such as overcomplicated ETL and redundant data staging. Although there might be good reasons for file staging (see the second blog above), in most cases I consider it a lunacy to stage perfect relational data to files, along the way losing metadata, complicating ETL, ending up serverless, and then reloading the same data into a relational database (ODS in this case).</p></blockquote>
<p>I’ve heard that the vendor justified the lake effort by empowering data scientists to do ML one day. I’d argue that if that day ever comes, the likelihood (pun not intended) of data scientists working directly on the source schema would be infinitely small since more than likely they would require the input datasets to be shaped in a different way which would probably require another ETL pipeline altogether.</p>
<h1>Better Data Staging</h1>
<p>I don’t subject my clients to excessive file staging. My file staging litmus test is what’s the source data format. If I can connect to a server and get in a tabular (relational) format, I stage it directly to a relational database (ODS or DW). However, if it’s provided as files (downloaded or pushed, reference data, or unstructured data), then obviously there is no other way. That’s why we have lakes.</p>
<p><img decoding="async" loading="lazy" class="alignnone wp-image-9215" src="https://prologika.com/wp-content/uploads/2024/06/a-diagram-of-a-computer-data-processing-descripti-4.jpeg" alt="A diagram of a computer data processing Description automatically generated" width="468" height="444" srcset="https://prologika.com/wp-content/uploads/2024/06/a-diagram-of-a-computer-data-processing-descripti-4.jpeg 946w, https://prologika.com/wp-content/uploads/2024/06/a-diagram-of-a-computer-data-processing-descripti-4-300x284.jpeg 300w, https://prologika.com/wp-content/uploads/2024/06/a-diagram-of-a-computer-data-processing-descripti-4-768x728.jpeg 768w, https://prologika.com/wp-content/uploads/2024/06/a-diagram-of-a-computer-data-processing-descripti-4-705x668.jpeg 705w, https://prologika.com/wp-content/uploads/2024/06/a-diagram-of-a-computer-data-processing-descripti-4-450x427.jpeg 450w" sizes="auto, (max-width: 468px) 100vw, 468px" /></p>
<p>Fast forward a few years, and your humble correspondent got hired to assess the damage and come up with a strategy. Data lakes won’t do it. Lakehouses and Delta Parquet (a poor attempt to recreate and replace relational databases) won’t do it. Fabric won’t do it and it’s too bad that Microsoft pushes Lakehouse while the main focus should have been on Fabric Data Warehouse, which unfortunately is not ready for prime time (fortunately, we have plenty of other options).</p>
<blockquote><p>What will do it? Going back to the basics and embracing the “<a href="https://learn.microsoft.com/en-us/power-bi/guidance/center-of-excellence-microsoft-business-intelligence-transformation">Discipline at the core, flexibility at edge</a>” ideology (kudos to Microsoft for publishing their lessons learned). From a technology standpoint, the critical pieces are EDW and organizational semantic models. If you don’t have these, I’m sorry but you are not modern yet. In fact, you aren’t even classic, considering that they have been around for long, long time.</p></blockquote>
<p><img decoding="async" loading="lazy" src="https://prologika.com/wp-content/uploads/2017/06/060417_1725_PrologikaNe2.png" alt="" /><br />
Teo Lachev<br />
Prologika, LLC | Making Sense of Data<br />
<a href="https://prologika.com/wp-content/uploads/2016/01/logo.png" rel="attachment wp-att-12"><img decoding="async" loading="lazy" class="alignnone size-full wp-image-12" src="https://prologika.com/wp-content/uploads/2016/01/logo.png" alt="logo" width="165" height="45" /></a></p>
<p>&nbsp;</p>
<p>&nbsp;</p>
]]></content:encoded>
					
					<wfw:commentRss>https://prologika.com/prologika-newsletter-summer-2024/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
	</channel>
</rss>
