<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Marcin Szymaniuk, Author at TantusData</title>
	<atom:link href="https://tantusdata.com/author/marcin-szymaniuk/feed/" rel="self" type="application/rss+xml" />
	<link>https://tantusdata.com</link>
	<description>That uncovers wisdom.</description>
	<lastBuildDate>Wed, 09 Apr 2025 13:19:21 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=6.7.1</generator>

<image>
	<url>https://tantusdata.com/app/uploads/2023/01/cropped-Favicon-32x32.png</url>
	<title>Marcin Szymaniuk, Author at TantusData</title>
	<link>https://tantusdata.com</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>Vendor lock-in when selecting a Cloud Data Platform architecture.</title>
		<link>https://tantusdata.com/insights/vendor-lock-in-when-selecting-a-cloud-data-platform-architecture/</link>
		
		<dc:creator><![CDATA[Marcin Szymaniuk]]></dc:creator>
		<pubDate>Wed, 09 Apr 2025 13:13:19 +0000</pubDate>
				<category><![CDATA[CloudMigration]]></category>
		<category><![CDATA[CloudStrategy]]></category>
		<category><![CDATA[DataWarehouse]]></category>
		<category><![CDATA[VendorLockIn]]></category>
		<guid isPermaLink="false">https://tantusdata.com/?post_type=insights&#038;p=2189</guid>

					<description><![CDATA[<p>Cloud Data Platform migrations come with hidden exit costs. Learn how to reduce vendor lock-in risk through smart architecture and technology choices.</p>
<p>The post <a href="https://tantusdata.com/insights/vendor-lock-in-when-selecting-a-cloud-data-platform-architecture/">Vendor lock-in when selecting a Cloud Data Platform architecture.</a> appeared first on <a href="https://tantusdata.com">TantusData</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<figure class="wp-block-image size-full"><img fetchpriority="high" decoding="async" width="960" height="540" src="https://tantusdata.com/app/uploads/2025/04/Grafiki-2.jpg" alt="" class="wp-image-2190" srcset="https://tantusdata.com/app/uploads/2025/04/Grafiki-2.jpg 960w, https://tantusdata.com/app/uploads/2025/04/Grafiki-2-300x169.jpg 300w, https://tantusdata.com/app/uploads/2025/04/Grafiki-2-768x432.jpg 768w" sizes="(max-width: 960px) 100vw, 960px" /></figure>



<p></p>



<h2 class="wp-block-heading">Migrating Data Platform, data warehouse or data lake?</h2>



<p>When deciding to move your data to the cloud, many people focus on costs, expected gains, easier maintenance, or simplified development. Sometimes, lower costs or easier maintenance drive the decision. However, it’s crucial to also consider the potential cost of exiting the cloud. What happens if, at some point, you decide you no longer want to be on a particular cloud platform? This could happen due to rising costs, new technology options, or even legal or political reasons.</p>



<p>Did you know that the exit-cost of your data platform from a Cloud might be a more expensive project than an original migration to the Cloud?<br><br>If you plan to migrate your data platform to the cloud, thinking about future migration costs now is a sign of responsible migration. Understand what will be the cost in terms of dollars, time, and effort if at some point you decide to exit that specific cloud vendor. This means recognizing that exit costs aren’t always obvious and that you’ll need to consider the various aspects of vendor lock-in.<br></p>



<h2 class="wp-block-heading">What exactly are the risks associated with vendor lock-in:</h2>



<ul class="wp-block-list">
<li><strong>Long term costs rise</strong>. If you don’t have an easy to move alternative you are at risk of becoming a hostage.</li>



<li><strong>Lack of flexibility</strong> &#8211; even if your solution is great now you always risk that it will not be developed in the future.&nbsp;</li>



<li><strong>High transfer fees</strong>. Most clouds are charging you if you transfer data from it (egress fee). That needs to be calculated when planning a migration from a specific cloud vendor.</li>



<li><strong>Contractual agreements</strong> &#8211; Vendor lock-in is not only about proprietary technology or data formats. What’s in the contract might be another trap which might become painful in the future.</li>



<li><strong>Lost optimization opportunities &#8211; </strong>there is no such thing as free lunch. Cloud solutions usually are easier to start with but your engineers might lose ability to do low level tuning if you even need that.</li>
</ul>



<p><mark style="background-color:rgba(0, 0, 0, 0)" class="has-inline-color has-black-color">When planning your cloud migration, don’t overlook the long-term exit costs &#8211; ask your provider about compatibility of specific technology with other tools in the market.&nbsp;</mark><br></p>



<h2 class="wp-block-heading">What can you do to mitigate the risk:</h2>



<ul class="wp-block-list">
<li>Evaluate the technology and alternatives – consider if there are open-source alternatives to your chosen data warehousing solution or if it&#8217;s compatible with other vendors</li>



<li>If you decide on proprietary technology make sure the cost of in and expected exit cost is justified by what you gain (usually easier, faster development)</li>



<li>Consider hybrid cloud. It’s more expensive but gives you more flexibility in the long run.</li>



<li>Consider favouring well-known standards and open source technologies and data formats. A good example is Kubernetes &#8211; it’s a technology which you can have on your own servers as well as with any major cloud providers.</li>
</ul>



<p>Considering moving to the cloud providers, ping me on Linkedin (<a href="https://www.linkedin.com/in/marcin-szymaniuk/">https://www.linkedin.com/in/marcin-szymaniuk/</a>) for specific calculations.</p>



<h2 class="wp-block-heading">Summary</h2>



<p>The data space is moving fast.&nbsp;</p>



<p>Always ask yourself what is the risk that within 5 years you have to migrate again.</p>



<p>Always ask yourself what will happen if you can’t use the selected data platform anymore. How long notice would you need to migrate to another solution?&nbsp;</p>



<p></p>
<p>The post <a href="https://tantusdata.com/insights/vendor-lock-in-when-selecting-a-cloud-data-platform-architecture/">Vendor lock-in when selecting a Cloud Data Platform architecture.</a> appeared first on <a href="https://tantusdata.com">TantusData</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>LLMs can make your business soar or sink, so tread carefully.</title>
		<link>https://tantusdata.com/insights/llms-can-make-your-business-soar-or-sink-so-tread-carefully/</link>
		
		<dc:creator><![CDATA[Marcin Szymaniuk]]></dc:creator>
		<pubDate>Tue, 05 Dec 2023 12:28:30 +0000</pubDate>
				<category><![CDATA[AI Implementation]]></category>
		<category><![CDATA[Business Strategy]]></category>
		<category><![CDATA[ChatGPT]]></category>
		<category><![CDATA[LLM]]></category>
		<category><![CDATA[ROI on AI]]></category>
		<guid isPermaLink="false">https://tantusdata.com/?post_type=insights&#038;p=1861</guid>

					<description><![CDATA[<p>Explore AI's impact on business; key steps for a strategic approach to LLMs, ensuring ROI &#038; privacy.</p>
<p>The post <a href="https://tantusdata.com/insights/llms-can-make-your-business-soar-or-sink-so-tread-carefully/">LLMs can make your business soar or sink, so tread carefully.</a> appeared first on <a href="https://tantusdata.com">TantusData</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<figure class="wp-block-image size-large"><img decoding="async" width="1024" height="585" src="https://tantusdata.com/app/uploads/2023/12/DALL·E-2023-12-05-13.27.17-Edit-the-second-image-of-a-businessman-at-a-crossroads-transforming-it-into-a-horizontal-rectangular-format.-In-this-revised-image-the-businessman-s-1024x585.png" alt="" class="wp-image-1875" srcset="https://tantusdata.com/app/uploads/2023/12/DALL·E-2023-12-05-13.27.17-Edit-the-second-image-of-a-businessman-at-a-crossroads-transforming-it-into-a-horizontal-rectangular-format.-In-this-revised-image-the-businessman-s-1024x585.png 1024w, https://tantusdata.com/app/uploads/2023/12/DALL·E-2023-12-05-13.27.17-Edit-the-second-image-of-a-businessman-at-a-crossroads-transforming-it-into-a-horizontal-rectangular-format.-In-this-revised-image-the-businessman-s-300x171.png 300w, https://tantusdata.com/app/uploads/2023/12/DALL·E-2023-12-05-13.27.17-Edit-the-second-image-of-a-businessman-at-a-crossroads-transforming-it-into-a-horizontal-rectangular-format.-In-this-revised-image-the-businessman-s-768x439.png 768w, https://tantusdata.com/app/uploads/2023/12/DALL·E-2023-12-05-13.27.17-Edit-the-second-image-of-a-businessman-at-a-crossroads-transforming-it-into-a-horizontal-rectangular-format.-In-this-revised-image-the-businessman-s-1536x878.png 1536w, https://tantusdata.com/app/uploads/2023/12/DALL·E-2023-12-05-13.27.17-Edit-the-second-image-of-a-businessman-at-a-crossroads-transforming-it-into-a-horizontal-rectangular-format.-In-this-revised-image-the-businessman-s.png 1792w" sizes="(max-width: 1024px) 100vw, 1024px" /></figure>



<h2 class="wp-block-heading">What should a Business Leader know before investing in AI?</h2>



<p>ChatGPT is everywhere. Chances are, you&#8217;re already using it for everyday tasks &#8211; its language skills are unmatched. Using it to fix spelling in your email is really easy, but let&#8217;s imagine transforming it into a frontline warrior – a chatbot that not only chats but also solves customer problems. We will use this business application of a chatbot as an example in this article.</p>



<p>Applying chatbots to more complex business processes is very tempting but can be full of pitfalls. There are many aspects to consider when building such an application. From the <strong>privacy</strong> of your data to the <strong>cost</strong> of implementation. And from the <strong>correctness</strong> of the responses provided by chat to the maintenance in the future. In this article, we will take all the challenges one by one, describe them and give a guideline on how to approach them pragmatically. This leads to building a successful project, providing a return on investment.</p>



<h3 class="wp-block-heading">What is a Large Language Model? What LLM is not?</h3>



<p>Remember the last time you chatted with ChatGPT? The conversation was smooth, and it even knew who the US president was. But pose a niche question about, say, <em>&#8216;how to change the invoice number in my system,&#8217;</em> and it might stumble. Even more, it might confidently tell you something off-base.&nbsp;</p>



<p>I bring that up because it&#8217;s essential to understand that ChatGPT or any other general <strong>LLM is not a general source of knowledge</strong>. Picture this: You wouldn&#8217;t ask a history professor about advanced rocket science, right? Similarly, out of the box, an LLM may not ace a deep dive into your specific business.</p>



<p>But it becomes a potent tool when you plug it with information about your domain and instruct it on what to do with that knowledge. A tool that might appear intelligent.</p>



<h2 class="wp-block-heading">Getting correct answers</h2>



<p>Imagine AI as a car and data as its fuel. Without the right fuel, it won&#8217;t take you where you want. Suppose you&#8217;ve set up your AI for customer support, and someone asks, &#8216;How do I get the invoice?&#8217;. Instead of the AI drawing a blank or guessing, it should know exactly where to fetch the answer &#8211; like a librarian who knows which shelf a book is on.</p>



<p><img decoding="async" width="624" height="349" src="https://lh7-us.googleusercontent.com/pzmgT0fCH2t5onAzeymAR4E-zJ1X1A5bxvuiOkUfidyVZHbaBHI2eu0kzQStT0q5ggoM5UWILLVRxw-gEPDJaKhPhwzuB7xbOldhdLFWHke7dnLj8YgzafJZ2G8jUUf87T0nTb5QjBW7BMWuR8uGGcw"></p>



<p>Data is the fuel for AI. For the chat to answer questions specific to your organisation, you need to provide it with the relevant information. Let&#8217;s say you want it to serve customer support purposes. Let&#8217;s assume the user asks the system,&nbsp;<em>&#8216;How do I get the invoice?&#8217;&nbsp;</em></p>



<p>The most standard approach to this problem is building an application that searches for relevant information and provides it to the chat. In our case, the application seeks information about the invoicing process. All searching is done in vector databases, which are great tools for finding the appropriate information in text documents. No wonder they&#8217;re becoming more popular since the rise of LLMs. Once it finds the clue, it hands it over to our chat.&nbsp;</p>



<p>Let&#8217;s stop here and amplify the message &#8211; general purpose LLM&nbsp;<strong>does not know much about your business</strong>. Here&#8217;s the gist: While AI is great with words, it only knows the specifics of your business if you tell it. Feed it the correct details to transform it from a general chat buddy to a helpful assistant. Think of it as training a new employee.</p>



<p>In the simple case described above, most of the &#8216;magic&#8217; is done by vector databases. But in some cases, just providing the information &#8216;on the fly&#8217; is not good enough. For the LLM to perform better, you must fine-tune the vector embeddings or the LLM model itself. This is more advanced, so we will describe it in a separate article.</p>



<h2 class="wp-block-heading">Privacy concerns</h2>



<p>Imagine sending a personal letter and having someone else read it. That happens when you use ChatGPT or similar tools &#8211; the data goes to their home base. No worries if you&#8217;re chatting about the weather. But what if it&#8217;s confidential customer details? That&#8217;s where things get tricky.</p>



<p>Picture a customer sharing personal info on your support chat. Now, where does that data go? Outside your walls? And are you breaking any rules by letting it?</p>



<p>Confirming the legal aspects is the minimum you should do. But if you are not allowed to send the data to third parties, you need to consider&nbsp;<strong>privately hosted LLM</strong>&nbsp;so your data never goes outside your data centre. The selection of open-source LLMs available to use as privately hosted is broad. Broad selection means you have many options, but at the same time, you need expertise to make wise selections. Especially since the market is hot and keeping up with all the changes is challenging. The right decision requires careful consideration of aspects like the model&#8217;s performance on your hardware and the ability to scale up and down if needed.</p>



<p>A cherry on top? Going private means you aren&#8217;t tied down by vendor lock-in. You can move your infrastructure to your own data centre or cloud. It&#8217;s a critical point in the context of mitigating the risk of a vendor making drastic changes in the pricing or in the offered service itself.</p>



<h2 class="wp-block-heading">Cost and defining the scope of the project</h2>



<p>When building a house, you wouldn&#8217;t start with the roof or fancy decor, right? You&#8217;d plan the foundation and budget from day one. It&#8217;s the same with an AI. If your supplier changes the rules, you&#8217;ve got expenses like engineering time, hardware, cloud, and possible surprise fees.</p>



<p></p>



<p><img loading="lazy" decoding="async" width="624" height="349" src="https://lh7-us.googleusercontent.com/rQZ1lQUbLCYOT2aC1DzMtMrZaNsYbKSUcKRCF3Cx6HLM4aFVhRFBVxtc9fkPAflAK2aO-41mdSjrXhSB1OTS89vSHPRtSAuLKGCJ1uNpuNqO_YMqpTGO8hQ65HAijjU1Eq7PwbI4_BSIUwq_BOxpv7E"></p>



<p>When building an AI application based on Language models, it&#8217;s essential to control the project&#8217;s scope and expected costs from day one.&nbsp;</p>



<p>It&#8217;s easy to dream big. After all, you&#8217;ve got all the tools at your fingertips. But to deliver the business value, you need to analyse the problem you are solving and set realistic milestones for the AI project. And it&#8217;s essential to start with issues which are common and easy to solve. E.g. When building a customer support application, it makes sense to begin by addressing issues which are repeated over and over again. You&#8217;d know them if you chat with your support team or skim through past tickets. If you approach the problem correctly, it&#8217;s likely that after two or three iterations with your product, you realise that you have solved 80% of the issues. Suddenly, you&#8217;ve got a solid ROI and can decide if you want to add those fancy trims or if your house is good enough.</p>



<h2 class="wp-block-heading">Maintenance</h2>



<p>Keep in mind that the more complex the application you are building, not only the implementation cost but also the cost of using it in the future increases &#8211; the maintenance and the price you pay for hardware or API calls. The maintenance might include aspects such as deployment complexity, testing, reacting to changes introduced by third parties or feeding your application with newer data.</p>



<p>That&#8217;s another reason to think early about the scope and be flexible during the implementation. So, before diving in, sketch out your costs. Ask yourself: How much is each chatbot chat worth to you? How much would you pay to help one customer if it&#8217;s a support bot? What conversion rate justifies the effort and cost of maintaining a chatbot if it&#8217;s for shopping? Knowing your numbers helps you steer clear of unwanted surprises.</p>



<h2 class="wp-block-heading">Summary</h2>



<p>As outlined above, it&#8217;s critical to understand the challenges when selecting the scope of the LLM project. But with a clear understanding of potential roadblocks from the beginning, you&#8217;re paving the way to the ROI you expect. The spectrum of challenges in implementing LLM for solving business problems in your organisation is vast, and you should:</p>



<ul class="wp-block-list">
<li>Ensuring&nbsp;<strong>correct</strong>&nbsp;behaviour makes it essential to provide your LLM with correct, well-structured data.</li>
</ul>



<ul class="wp-block-list">
<li>Addressing&nbsp;<strong>privacy</strong>&nbsp;concerns means understanding the legal prerequisites. Remember that you have options of hosting LLM in your Data Center, which simplifies the legal aspects but introduces extra technical challenges.</li>
</ul>



<ul class="wp-block-list">
<li>Budgeting for a project means considering all&nbsp;<strong>costs</strong>: development, API charges, and maintenance. It&#8217;s important to verify the assumptions after every milestone of the project and make sure the scope and the cost are under control.</li>
</ul>



<ul class="wp-block-list">
<li>Thinking ahead is essential, so estimate&nbsp;<strong>maintenance</strong>&nbsp;efforts. Ensure these efforts align with the expected benefits of its integration into your organisation.</li>
</ul>
<p>The post <a href="https://tantusdata.com/insights/llms-can-make-your-business-soar-or-sink-so-tread-carefully/">LLMs can make your business soar or sink, so tread carefully.</a> appeared first on <a href="https://tantusdata.com">TantusData</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>GDPR, the forgotten done right</title>
		<link>https://tantusdata.com/insights/gdpr-the-forgotten-done-right/</link>
		
		<dc:creator><![CDATA[Marcin Szymaniuk]]></dc:creator>
		<pubDate>Mon, 03 Apr 2023 06:00:00 +0000</pubDate>
				<category><![CDATA[anonymisation]]></category>
		<category><![CDATA[data erasure]]></category>
		<category><![CDATA[GDPR]]></category>
		<category><![CDATA[right to be forgotten]]></category>
		<guid isPermaLink="false">https://tantusdata.com/?post_type=insights&#038;p=1439</guid>

					<description><![CDATA[<p>In the previous article, we covered anonymization and pseudonymization &#8211; techniques used in the context of ensuring data privacy, and more specifically, in the context of GDPR. While the basic idea and goal of anonymization are intuitive, GDPR introduces more subtle regulations, such as the Right to Be Forgotten. This concept is tricky because it [&#8230;]</p>
<p>The post <a href="https://tantusdata.com/insights/gdpr-the-forgotten-done-right/">GDPR, the forgotten done right</a> appeared first on <a href="https://tantusdata.com">TantusData</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="576" src="https://tantusdata.com/app/uploads/2023/03/watching-you-1024x576.jpg" alt="Big data and GDPR, user's right to be forgotten" class="wp-image-1440" srcset="https://tantusdata.com/app/uploads/2023/03/watching-you-1024x576.jpg 1024w, https://tantusdata.com/app/uploads/2023/03/watching-you-300x169.jpg 300w, https://tantusdata.com/app/uploads/2023/03/watching-you-768x432.jpg 768w, https://tantusdata.com/app/uploads/2023/03/watching-you-1536x863.jpg 1536w, https://tantusdata.com/app/uploads/2023/03/watching-you-2048x1151.jpg 2048w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></figure>



<p style="font-size:18px">In the previous article, we covered anonymization and pseudonymization &#8211; techniques used in the context of ensuring data privacy, and more specifically, in the context of GDPR. While the basic idea and goal of anonymization are intuitive, GDPR introduces more subtle regulations, such as the Right to Be Forgotten. This concept is tricky because it is subject to interpretation. The complexity is further compounded by technical implementation details within a complex ecosystem that handles large datasets.</p>



<h2 class="wp-block-heading"><strong>Open to interpretation</strong></h2>



<p style="font-size:18px">If you possess customer data and the customer revokes their consent, requesting to be forgotten, you must take action. You need to ensure that you remove the data. What&#8217;s tricky about this? There are at least two aspects to consider:</p>



<p>1. If you decide to simply delete the data, ensure that the deletion actually occurs on a physical level. Modern data platforms excel at abstracting certain operations, so just because you think you deleted the data doesn&#8217;t mean it&#8217;s unrecoverable or that it&#8217;s no longer present in your systems.</p>



<p>2. Alternatively, you might not need to remove the data if you choose to anonymize it instead. In this case, you must ensure that you correctly identify all the Personally Identifiable Information (PII) fields and that the anonymization process is accurate. More on this topic follows below.</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="576" src="https://tantusdata.com/app/uploads/2023/03/detail_abstract-1024x576.jpg" alt="The details count, you have to pay attention to the deletion behaviour in each tool you use." class="wp-image-1444" srcset="https://tantusdata.com/app/uploads/2023/03/detail_abstract-1024x576.jpg 1024w, https://tantusdata.com/app/uploads/2023/03/detail_abstract-300x169.jpg 300w, https://tantusdata.com/app/uploads/2023/03/detail_abstract-768x432.jpg 768w, https://tantusdata.com/app/uploads/2023/03/detail_abstract-1536x863.jpg 1536w, https://tantusdata.com/app/uploads/2023/03/detail_abstract-2048x1151.jpg 2048w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></figure>



<h2 class="wp-block-heading"><strong>The devil lies in the implementation</strong></h2>



<p style="font-size:18px">Deleting data when a customer revokes their consent may seem appealing from a security perspective. The assumption is that once the data is gone, you are safe, right? While this is true, the deletion process itself is not so straightforward. The complexity arises from the varied ways in which tools within the Data &amp; Analytics ecosystem handle data removal. The deletion behavior in these tools does not always align with the general intuition of &#8216;deletion.&#8217;</p>



<p style="font-size:18px">Most modern data processing tools are optimized for ingestion and analytics queries, and their ability to perform fine-grained data deletion is limited. Sometimes, this means that you cannot simply remove individual records. At other times, it means that the deletion is not immediate. Finally, it might be that the data is merely marked as unavailable, but not actually deleted.</p>



<p style="font-size:18px">Do you see how some of these options might not be acceptable from a regulatory perspective? Therefore, it is crucial to consider technical limitations when deciding on a strategy for the Right to Be Forgotten.</p>



<p>Below, we will explore some commonly used data processing tools:</p>



<p>1. Traditional relational databases &#8211; These databases are good at operations involving single rows at a time, making it easy and safe to delete specific rows.</p>



<p>2. Hadoop and HDFS &#8211; It is impossible to delete a single record. If you want to remove the data of a single user, you need to rewrite the entire dataset. The rewrite comes with its own considerations, but the main one is whether it&#8217;s acceptable from a performance and cost perspective to rewrite the entire dataset just to remove a single record.</p>



<p>3. Delta, Hudi, Iceberg &#8211; These systems allow for the deletion of individual records, but you need to understand that the &#8216;deletion&#8217; is merely creating another delete-marker record. If you want to ensure that the data physically disappears eventually, you need to carefully design the compaction strategy specific to the storage you are using.</p>



<p>4. NoSQL databases &#8211; Each NoSQL database has its own way of handling deletion. For instance, Cassandra uses tombstones &#8211; data-deletion markers &#8211; which lead to similar concerns as described in the previous point. A carefully designed compaction strategy is a must.</p>



<p>5. BigQuery &#8211; Large-scale deletion is not what BigQuery is optimized for, which is why the tool has quotas for these kinds of operations. You also need to be aware that the data does not disappear immediately. That&#8217;s why it&#8217;s worth considering a full rewrite or crypto shredding (please refer to the next point) in certain scenarios.</p>



<p><em>Does this cover all options? </em></p>



<p>Unfortunately, the variety of tools makes it impossible to mention them all in a single article. You have to carefully investigate your ecosystem in the context of the Right to Be Forgotten. I only mention some of the commonly used tools to illustrate the complexity of the problem.</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="576" src="https://tantusdata.com/app/uploads/2023/03/left-behind-1024x576.jpg" alt="Anonymised data doe snot equal forgotten. The physical evidence remains like a forgotten luggage anyone can look through to de-anonymise the data again." class="wp-image-1442" srcset="https://tantusdata.com/app/uploads/2023/03/left-behind-1024x576.jpg 1024w, https://tantusdata.com/app/uploads/2023/03/left-behind-300x169.jpg 300w, https://tantusdata.com/app/uploads/2023/03/left-behind-768x432.jpg 768w, https://tantusdata.com/app/uploads/2023/03/left-behind-1536x863.jpg 1536w, https://tantusdata.com/app/uploads/2023/03/left-behind-2048x1151.jpg 2048w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></figure>



<p>The right to be forgotten is not fulfilled if the data is not erased on a physical level. Then it&#8217;s simply hidden. Left behind till somebody finds it.</p>



<h2 class="wp-block-heading"><strong>Does anonymised = forgotten?</strong></h2>



<p>Deleting records is not the only option for forgetting a customer. You can consider anonymizing the records of customers who revoke their consent. Anonymizing the dataset has the obvious benefit of retaining some of the data (without PII) so it can be used for analytics purposes in the future.</p>



<p>However, remember that you need to be very careful when designing the architecture and procedures. Things to keep in mind:</p>



<p>1. Anonymization should be a non-reversible process.</p>



<p>2. You have to correctly identify and anonymize the PII fields.</p>



<p>3. Anonymizing just PII fields might not be enough. You have to ensure that you prevent indirect identification of the user&#8217;s data. More about this will be covered in the next article.</p>



<p>4. You must have a formal process in place and follow it when handling your datasets. Automate as much as possible to limit the chance of human errors and privacy breaches.</p>



<p>5. Naive anonymization of individual records comes with problems similar to deletion itself. You have to consider the properties of the analytics tools in your stack before deciding on this approach.</p>



<p>One of the specialized techniques for handling the Right to Be Forgotten is crypto-shredding. The basic idea is to encrypt customer records and remove the encryption key when the customer requests to be forgotten. Using this technique can free you from concerns related to the cost of reprocessing the entire dataset. At the same time, it must be carefully thought through, especially when it comes to storing, securing, and removing the encryption keys.</p>



<p>Look for a description of crypto-shredding in one of the upcoming articles.</p>



<h2 class="wp-block-heading"><strong>What Have You Missed?</strong></h2>



<p>Once you have determined the appropriate method for addressing the Right to Be Forgotten (RTBF), you are nearly prepared to begin the implementation process. Why nearly? A crucial aspect of this process is identifying all datasets containing sensitive information. It goes without saying that managing your master data is essential. However, are you fully aware of the life cycles of your datasets? Can you confidently say that you are handling all derived datasets? In other words, are you dealing with a data lake, data lakehouse, or perhaps a data swamp?</p>



<p>Maintaining control over data lineage is a critical component of Data Governance, and we will discuss its significance in the context of RTBF in an upcoming article.</p>
<p>The post <a href="https://tantusdata.com/insights/gdpr-the-forgotten-done-right/">GDPR, the forgotten done right</a> appeared first on <a href="https://tantusdata.com">TantusData</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Obtaining value from GDPR with solutions that work for your bottom line.</title>
		<link>https://tantusdata.com/insights/obtaining-value-from-gdpr/</link>
		
		<dc:creator><![CDATA[Marcin Szymaniuk]]></dc:creator>
		<pubDate>Fri, 27 Jan 2023 11:02:39 +0000</pubDate>
				<category><![CDATA[anonymisation]]></category>
		<category><![CDATA[compliance]]></category>
		<category><![CDATA[GDPR]]></category>
		<category><![CDATA[pseudonymisation]]></category>
		<guid isPermaLink="false">http://tantusdata.local/?post_type=insights&#038;p=571</guid>

					<description><![CDATA[<p>Compliance with the GDPR regulations can be profitable when done right. Apart from saving on legal fees and avoiding customer attrition, you can also prevent issues and support your B2B data sales pitches. For that you need to cover all aspects which we look into below. GDPR regulations have certainly been around for quite a [&#8230;]</p>
<p>The post <a href="https://tantusdata.com/insights/obtaining-value-from-gdpr/">Obtaining value from GDPR with solutions that work for your bottom line.</a> appeared first on <a href="https://tantusdata.com">TantusData</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="576" src="https://tantusdata.com/app/uploads/2023/01/data-1024x576.jpg" alt="big_data_gdpr" class="wp-image-1496" srcset="https://tantusdata.com/app/uploads/2023/01/data-1024x576.jpg 1024w, https://tantusdata.com/app/uploads/2023/01/data-300x169.jpg 300w, https://tantusdata.com/app/uploads/2023/01/data-768x432.jpg 768w, https://tantusdata.com/app/uploads/2023/01/data-1536x863.jpg 1536w, https://tantusdata.com/app/uploads/2023/01/data-2048x1151.jpg 2048w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></figure>



<p style="font-size:18px">Compliance with the GDPR regulations can be profitable when done right. Apart from saving on legal fees and avoiding customer attrition, you can also prevent issues and support your B2B data sales pitches. For that you need to cover all aspects which we look into below.</p>



<p style="font-size:18px">GDPR regulations have certainly been around for quite a while already. Naturally everybody in the business world is familiar with the basic ideas behind it and companies have implemented their compliance solutions. Is it sufficient? Very often all the business and technical aspects are not part of these. Have you missed any? Did you consider all those that have profit potential? Selected of the often neglected details will be mentioned below, others will be covered in next articles. You can consider this article as a MVP checklist.</p>



<p>Circling back to that profit – and not simply in terms of avoiding hefty fines, but those of added value. Having good processes in place along with assurance of security and speed creates competitive advantage for those who work with and share data with 3rd party companies. Ability to indisputably prove to your partners that your data is well handled and by using it they are not facing any GDPR breach risks serves as an additional bonus. One that is likely to shift the negotiations in your favour, as few companies bother to follow in-depth GDPR compliance recommendations and stick to basic requirements. You stand out for the win.&nbsp;</p>



<p>If you are using an external company to verify and enhance your GDPR-compliant solutions, then you can ask them to document how well your company is performing in this aspect (including the solutions they introduced). This makes it easier to verify the quality for your partners.</p>



<h2 class="wp-block-heading">Keep in mind: analytics and security departments have contradicting goals</h2>



<p>Whilst your analytics department is happy to keep as much data as possible, from the security team’s point of view: less is more. The less data, the more safety. Even if GDPR rules are in place, they will opt for minimising the data amount for a lower breach and/or leak risk. On the other hand the value data contain is substantial and you would miss out if you strictly followed the security department’s preference. Therefore you must strategise to reach the right balance between opportunity and risk. There are two main techniques associated with GDPR data and these are already impacting this balance.</p>



<h2 class="wp-block-heading">Pseudonymisation (pseudo anonymisation) </h2>



<p><strong>Pseudonymisation</strong> is a process which can be reverted with use of extra information, such as e.g. salt, encryption key. Using that information you can calculate which customer the data is associated with the pseudonymised data. It is a very fine method and many companies leverage it in some way. However, the key question is whether it is done right. You have to keep in mind a few key elements.</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="325" src="https://tantusdata.com/app/uploads/2023/01/GDPR_data_pesudonymisation_TantusData-1024x325.jpg" alt="What to check to see if pseudonymisation is correct:
right purpose, consent from users, extra information is securely stored, PII correctly identified, reverting process impossible without the key" class="wp-image-1366" srcset="https://tantusdata.com/app/uploads/2023/01/GDPR_data_pesudonymisation_TantusData-1024x325.jpg 1024w, https://tantusdata.com/app/uploads/2023/01/GDPR_data_pesudonymisation_TantusData-300x95.jpg 300w, https://tantusdata.com/app/uploads/2023/01/GDPR_data_pesudonymisation_TantusData-768x244.jpg 768w, https://tantusdata.com/app/uploads/2023/01/GDPR_data_pesudonymisation_TantusData-1536x487.jpg 1536w, https://tantusdata.com/app/uploads/2023/01/GDPR_data_pesudonymisation_TantusData-2048x649.jpg 2048w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></figure>



<p>Bear in mind that it is very easy to identify fields which are obviously sensitive such as the name or a personal number. This often causes an illusion that PII seems an easy variable to spot. And it results in many companies missing those that do not seem sensitive at first glance. Such include those attributes that can be used to identify a single user when used in combination with other variables. More about that later on.</p>



<h2 class="wp-block-heading">Pseudonymisation: pros &amp; cons</h2>



<p>Why choose pseudonymisation? Obviously you would like to be able to allow data analysts and data scientists to work with the data without exposing the personal information. This technique enables data scientists to e.g. come up with an upsell offer without seeing the personal information, thus meeting GDPR requirements. The key advantage of this approach is that it is reversible. This can be leveraged during offer send out, which at this particular step required the knowledge of which tailor made offer is best for which customer. Additionally this step can be fully automated. Subsequently, there is no need for data scientists to engage with the personal data.<br>&nbsp;The fact that <em>pseudonymization</em> is reversible can be used during offer send out &#8211; only at that step there is a need to know which customer to send the specific offer. And this step can be fully automated so there is no need of touching personal data by data scientists</p>



<h2 class="wp-block-heading">Anonymisation:</h2>



<p><strong>Anonymisation</strong> is a process which is not reversible. This means that when working with anonymised data, one should never be able to identify an individual person based on the data. Yet again, there are a few aspects one must keep in mind when implementing it and during the approach selection process.</p>



<p>On one hand, anonymisation when done properly is significantly safer as a method. So it permits doing more with a company’s data. Even selling it to 3rd parties. However, it does limit your ability as a B2C provider. You will be unable to get insights and actions which are dedicated to an individual customer. Both of these hinge on correct implementation.</p>



<p><em>PII must be correctly identified and excluded, so that the anonymisation is not reversible.</em> &#8211; Marcin Szymaniuk</p>



<p>The anonymisation should be 100% irreversible. Therefore, it is critical to properly analyse and identify all sets of attributes which can be used to identify a singular user. These must not be used. The reverse engineering could potentially happen after collecting an array of data points for many days. That&#8217;s why it&#8217;s important to go through careful analysis of all the data, including that which doesn’t belong to the sensitive field. For instance tracking the same customer for a long period of time could reveal his/her identity. The right analysis takes preemptive measures and will show all those that must be disallowed to present that. Disabling reverse engineering is particularly important for 3rd party players.</p>



<h2 class="wp-block-heading">Anonymisation: advantages &amp; drawbacks</h2>



<p>Data presents a significant value. This is true for both internal use and for external partnerships. Business customers that purchase data want to safeguard against potential GDPR breach implications. As there is now a high competition amongst owners of big data, who wish to profit from sharing, it is important to stand out. 3rd party customers in particular can have more restrictive criteria. With truly ensured safety from breaches and impossible reverse engineering you become a strong competitor. And you are likely to make your security department more lenient in terms of amounts of data you get to keep, as this is a more secure method. Just as long as the implementation is done correctly.</p>



<p>A significant limitation of this method is the inability to target individual customers in your analysis and highly targeted personalised offers. Though it can be very useful for internal purposes as well. Such data still allows understanding trends and collecting insights about specific demographics and so on.&nbsp;</p>



<h2 class="wp-block-heading">Final comments</h2>



<p>GDPR compliance is a broad subject. Particularly when it comes to optimising the solutions for added value. This article has briefly covered the importance of analysis, particularly when it comes to variables that in combination with others enable identifying a single individual. There are multiple techniques, only two mentioned. More details and breadth will be covered in the upcoming articles. You can also expect subjects such as the multiple ways to approach the ‘right to be forgotten’ and ‘right to access’. So stay tuned.</p>



<p>Follow us on <a href="https://www.linkedin.com/company/tantusdata/">LinkedIn</a></p>
<p>The post <a href="https://tantusdata.com/insights/obtaining-value-from-gdpr/">Obtaining value from GDPR with solutions that work for your bottom line.</a> appeared first on <a href="https://tantusdata.com">TantusData</a>.</p>
]]></content:encoded>
					
		
		
			</item>
	</channel>
</rss>
