Sponsored Archives - Gestalt IT: Recent Episodes

None

The Latest News in Enterprise IT, part of The Futurum Group

View Details

There comes a time in certain relationships where you have to bring them to an end. It could be they’re a drain on resources, where you have to give more than you get. Or it could be that you don’t feel like you own anything anymore. Yeah, I’m talking about working with Cloud Service Providers (CSPs).

The cost of services can exceed the value they provide for hosting, and the ability to schedule tasks on your timeline can be exorbitant. The data you gather and generate is held hostage for a monthly fee, even if it seems insignificant to start with.

Sometimes it really is necessary to say, “It isn’t me, it’s you.” And say goodbye to cloud providers.

Repatriating, or importing data from a cloud provider to a private data center is never an easy task. Building a data center is daunting – you need to get the space, the power, and the bandwidth yourself to make it work.

A Friend to Help You MoveJust like moving yourself or your family, there are some tasks that can be exciting and fun. Figuring out a data center architecture or the provider agreements can be interesting negotiations, just like choosing a new place and finding furnishings.

And then the time comes to move your “stuff.” The boring and tedious part. You can be someone who is okay with scuffed tables, broken kitchenware, and scrambled belongings. Then again, there’s a time to call in professionals.

Interlock Technology is a team of professionals that can help move your data from service providers to your private cloud. Interlock offers what it calls A3 Migration Solutions to facilitate and ease data migration. I participated in a showcase with Noemi Greyzdorf, VP of Operations, and Massimo Yezzi, CTO.

A3 (say “A cubed”) stands for “Any Data, Any Storage, Anywhere,” which is the basis of their solution. Their migration capabilities allow any object and file data to be migrated in an application-aware way. This means the data maintains consonance (if you have to look it up: in agreement and compatible between actions), compliance, permissions and chain of custody. All of these attributes are preserved end to end.

A perfect use case is repatriating application data from a public service provider to your private cloud. The application performance from the customer’s perspective is seamless and transparent.

ScreenshotInterlock’s A3 solutions are offered in two different ways. A3 DF Classic is a managed, professional services engagement with personnel from Interlock ensuring complete and successful data migration. Alternatively, there’s A3 DATAFORGE, a licensed software solution used by a certified migration engineer.

Consulting is available from the Interlock team, but the engagement isn’t a full professional services endeavor.

Fast, Easy DeploymentUnder the hood, the A3 solutions have a couple different components: the Data Application Manager and Data Movers. The Data Application Manager (ForgeManager) is a single instance that oversees the migration and provides the migration engineer a dashboard of progress and status. Multiple Data Movers facilitate the migration, each capable of supplying up to 128 threads of migration. Data Movers are deployed as needed to accomplish the migration at the speed the customer is looking for.

The ForgeManager and the Data Movers are simply VMs that can be deployed into a customer’s existing VMware virtual infrastructure. Interlock supplies them as OVFs that are quickly and easily installed onto existing ESXi instances.

ScreenshotTo pull application data from a source like AWS S3 storage to a target like a local NFS filer, just point the A3 solution at the source and target, connect the correct application-aware modules, and it’s good to go.

Repatriation of cloud data isn’t its only trick. Interlock is just as capable of migrating between disparate storage systems within a single data center or across long distances.

Consistent, Certifiable ResultsInterlock takes pains to ensure that migrated data is exactly the way an application needs it. Metadata, permissions, audit trails and logs are all retained. This is accomplished by using individual, application-aware modules. These can be off-the-shelf applications, or new modules can be developed for custom applications. The available library includes more than 20 commonly-used commercial applications, like NICE, Documentum and SourceOne.

Regulatory compliance is maintained, adhering to stringent requirements concerning data integrity, immutability, chain of custody, authentication and permissions. All ownership and retention information associated with the data is migrated along with the data. Encrypted data isn’t decrypted, and multiple checksums, like MD5 and SHA256, ensure the integrity of the migration.

Finally, audit reports detail each action performed during the migration, guaranteeing that the full data lifecycle is documented.

DATAFORGE migrates data at the storage layer rather than the application layer to speed up the process. This enables large-scale datasets to be migrated in a fraction of the time they would otherwise take. This is useful when data is trapped because an application can no longer access the data where it resides. This could be the case if the license on a particular version of an application expires, or an outdated storage system starts to fail and can’t handle an application workload.

Scalability and ObservabilityInterlock’s A3 solutions provide a framework to easily scale out to numerous migration tasks associated with various applications. It can also scale up to accommodate large-scale applications involving billions of objects and many terabytes of data.

This can all be observed through their ForgeManager as individual tasks with various different states and progress.

Screenshot
Each task can be expanded to view individual status. Errors can be checked and remediated. Simplicity is one of the goals of the A3 Solution.

ConclusionIn the case of repatriation, the data where it resides is simply costing money, which could be saved in an on-premises data center. Repatriation can also be needed when an application can no longer run on a cloud service provider, either due to enterprise guidelines or regulatory reasons.

Private cloud is the way forward for many organizations. But like it or not, data migration is part of that. The A3 DATAFORGE solution by Interlock Technology provides a consistent, compliant and fast way of leaving public cloud, and move applications and data to the safety of an on-premises data center where you control them.

I’ve written and talked extensively about building private cloud infrastructure in a variety of places where you can learn how to settle into your own data center easier and faster. Talk to Interlock Technology about their case studies to get more info. And make sure to check out more coverage from my fellow panelists on the Tech Field Day Showcase with Interlock.


© Gestalt IT, LLC for Gestalt IT: Leaving Cloud Ain’t Easy: A Roadmap to Repatriate Data with Interlock

View Details

As an old-school database administrator (DBA), I’ve overseen more than my share of migrations between database platforms, both on-premises and in the cloud. The projects required considerable planning, experimentation, and zero tolerance for data loss.

Fortunately, they were all successful because modern databases have robust backup and recovery mechanisms in addition to sophisticated transaction logging to guarantee data synchronicity during transference.

Databases depend on block storage for data retention, and that moreover enables DBAs to easily ascertain accidental corruptions, if any, within a single block.

But when dealing with data that’s read and written at the application level – often stored within thousands, even millions of individual files and retained within object storage – migrating entire applications without losing any data while keeping them executing is a much trickier challenge. The process can consume days or even weeks as files containing terabytes of critical application data are ported between source and target environments.

Consonance, Security, and Chain of CustodyAny migration process must guarantee application consonance, in other words, seamless execution while the app’s underlying data resides at the source or target. Consonance must be maintained regardless of the file’s current location – be it a local file storage or a cloud “bucket”, and the access method, whether using REST APIs or application-specific protocols.

But consonance is only the first concern. Data that’s encrypted for security purposes must remain encrypted during transfers. In cases of extremely sensitive data, it must be seen that only the migration tool touches it, so that a legal chain of custody is maintained.

Even more problematic is that it can take considerable computing power and reliable networking to migrate application files at speed, and that could mean compromising current hardware and bandwidth at the source right in the middle of crucial business cycles.

DATAFORGE: A Migration Solution, CubedI recently came to know about a comprehensive solution that squarely addresses these data migration challenges. From Noemi Greyzdorf (VP ofOperations) and Massimo Yezzi (CTO) of Interlock, I heard about DATAFORGE, a solution designed to satisfy the potentially tricky and voluminous object migration requirements.

Interlock calls DATAFORGE their A-cubed (A3) solution for object migration. Designed to move any data between any storage, anywhere, it guarantees application consonance while handling all the other aforementioned issues.

ScreenshotDATAFORGE currently comes in two flavors:

  • For IT teams that are overwhelmed with work, Interlock offers DATAFORGE Classic. DATAFORGE Classic comes with Interlock’s team of experienced and certified migration engineers who handle the solution deployment and migration process end to end for enterprises.
  • For organizations that can spare resources, Interlock allows members of the staff to train to be DATAFORGE-certified migration engineers and manage the migration in-house.

Nuts and BoltsAt its heart, DATAFORGE uses a series of Data Movers – usually a VMware virtual machines configured via ESXi – and managed by a single DATAFORGE application node. As the name implies, Data Movers move application files from source to target, but what makes them unique is Interlock’s proprietary software.

As Data Movers continually copy data to the target system, the apps continue to communicate with the files stored at the source until cutover is completed at the target, thus ensuring application consonance. DATAFORGE remains outside the migration path, and since the Data Movers migrate data at the storage level only, this keeps them focused on completing the process as rapidly as possible.

DATAFORGE provides a web-based UI that constantly monitors the migration progress, and includes the ability to migrate files via multi-threaded migration processes.

Migration engineers can use DATAFORGE’s built-in performance scheduler to control the number of threads each Data Mover is using allowing them to throttle migration during peak application business hours, and ramp up migration processing during off-peak periods.

Migration is not just about moving files. It’s also crucial to maintain individual file’s metadata for handling security and retention requirements post-migration. Data Mover processes can be tuned to accommodate files whose contents rarely change – long-archived logs, documents, or images – that need to be retained historically for regulatory purposes. Data Movers interrogate those files’ metadata less frequently, and focus more efforts on volatile ones.

ScreenshotMaintaining Consonance Regardless of FormatMaintaining application consonance during migration is trickier than it appears. For one, the application may use several different connection methods to read from or write data to files. They may be using a vendor-defined REST API or even content-addressable storage (CAS) if there are specific security requirements to access and write data via cryptographic means.

Data within the files themselves are likely to be stored in any one of a plethora of proprietary formats: images (JPEG, JIFF, PNG), audio-visual recordings (MP4, M4A), geographic information systems (GeoJSON, GML, KML), standard documents (JSON, XML), and so on.

DATAFORGE’s software is designed to handle this seamlessly regardless of the file format. If required, Interlock can even build a custom application connector so that it appears to the application as though it had written data directly to the target storage system.

Never Break the ChainMaintaining the chain of custody for any secured data stored within the object file system is a crucial requisite. An encrypted file must remain so during the transfer process without the need to first decrypt it. Otherwise, data might be corrupted, to say nothing of the risks of security regulation violation through accidental glancing of the data.

DATAFORGE’s software calculates a hash on every bit comprising the file as the transfer proceeds, meaning it can guarantee that the file and its metadata is never modified while migrating. Interlock guarantees that its chain of custody reporting is solid enough to stand up in a court of law.

So How Fast is Fast?Interlock says that as long as there are adequate resources – essentially, a sufficient number of Data Movers and a robust network connectivity between source and target environments – a migration engineer should be able to move over 20 terabytes of files per day, and up to a billion files per week. That’s a pretty bold claim, but not one that is unsupported by successful customer stories:

  • The closing of an on-premises datacenter necessitated a complete migration to AWS S3 buckets in the cloud. This required DATAFORGE to migrate 176 million files over a 14-day period at a daily transfer rate of 24 terabytes. This case was rather unique, as the data was essentially trapped within the customer’s application but needed to be retained for 10 years to satisfy regulatory requirements. DATAFORGE satisfied the need to reliably migrate file metadata ensuring that individual retention periods for each file is honored at the target system.
  • For a very different customer use case, a financial institution specifically, a DATAFORGE configuration of eight multi-threaded Data Movers leveraged a 1GbE network connection to copy over 220 terabytes of data from the source system to its new target – an on-prem object store. DATAFORGE maintained application consonance during the 30-day period during which over three billion files with an average size of 75KB were successfully migrated across multiple protocols. DATAFORGE also produced a chain-of-custody report to satisfy regulatory requirements.

Conclusion: In closing, though my object storage migration experience is admittedly limited, I believe that Interlock’s DATAFORGE solution appears to check all the boxes I’d be concerned about when tasked with a cross-platform migration project.

For more, head over to the Tech Field Day website or YouTube channel to see videos from the Interlock Tech Field Day Showcase.


© Gestalt IT, LLC for Gestalt IT: Preserving Application Consonance while Reliably Migrating Files and Data with Interlock DATAFORGE

View Details

In IT, one thing you learn very quickly is that the only constant in this career path is change. Things in an environment are always moving and changing – including critical data. Cloud and hybrid-cloud transitions have caused legacy and application data to be shuttled between on-premise datacenters and the cloud, and now, back to the datacenter all over again.

ScreenshotIntroduction to DATAFORGEInterlock’s solution comes in two flavors – A3 DATAFORGE and A3 DF Classic. Released last year, A3 DATAFORGE is a software solution that allows scalable, managed migration between protocols and APIs like SMB, NFS, S3 or REST. It is designed to allow for scale up and scale out deployments that ensure both data integrity and regulatory compliance needs are met when the final cutover happens.

The platform supports migration of Any data to Any storage to Any Destination (A3), says Interlock.

ScreenshotDF Classic is a fully-managed solution that supports cross-protocol migrations using predeveloped connectors and custom versions as needed. It includes support for migration engineers to setup, monitor and cutover end-state data for a truly end-to-end solution.

The latter is a more SaaS option which can be licensed on a storage transferred basis and after initial assessment, is available to use by organizations’ own staff.

In either options, A3 DATAFORGE relies on 2 base components – DATAFORGE Manager and application nodes, that must be deployed close to the source or destination storage for reduced latency.

DATAFORGE Manager is a deployable OVA for VMware environments that acts as a management overlay for the entire system. The nodes are the actual data movers which can be package installed on any supported variant of Microsoft Windows in either a scale up or scale out model. Users can right-size the data mover throughput by either scaling up, making each VM larger in the number of cores assigned, or by scaling out and having multiple nodes with each capable of 128 threads, depending on the use case.

Regardless of what you choose, Interlock will assess what is to be moved and provide guidance on sizing of the solution.

ScreenshotWhen performing assessment, Interlock’s migration engineers focus not only on size and number of files but also the type of data. This assessment allows users to be more cost-efficient and prioritized in their migrations as they can identify data that may not need moving at all vs. that which must be prioritized.

For services such as Content Addressable Storage (CAS), Interlock can leverage its collector code to support application-aware migration.

ConclusionData migration is a service many companies are competing to provide. A very large segment of the IT world requires it to for their on-going hybrid multi-cloud journey. Unfortunately, migration can often be bulky, cumbersome and prolonged taking weeks to years on end. The processes are riddled with data loss scenarios. By creating a framework service that can be consumed in both a managed and self-service fashion, Interlock addresses many of these challenges of legacy migration enabling seamless data mobility and access of assets across platforms.

For more, head over to the Tech Field Day website or YouTube channel to see videos from the Interlock Tech Field Day Showcase.


© Gestalt IT, LLC for Gestalt IT: Legacy Data Migration is a Problem Worth Solving with Interlock

View Details

In the latest Tech Field Day Showcase, I alongside a group of other delegates, met with Interlock Technology. Interlock is a company that specializes in supporting highly complex data migration projects with a particular focus on critical and compliance data.

Data consistency and integrity are two absolute precursors in any data storage architecture stack, and were among the key differentiators pitched during the discussion. So why is this relevant in the context of data migration? Let me explain.

There are several ways to look at data consistency and integrity through the various layers of IT infrastructure stack. At the lowest layer, you want to ensure that integrity is embedded into the system by making certain that the endless chains of ones and zeroes that constitute a file or object are properly written, readable, and protected against common hardware failures.

As you move up the stack, data integrity is essential to ensuring that data hasn’t been tampered with, notably when it comes to regulated data. This is the case for any WORM type data.

Nearly every IT professional must deal with data migration activities at some point in their career. Most times, these are small to moderate scale activities that are carried over through system utilities such as rsync or robocopy. Preserving state is paramount in these operations. At the bare minimum, you want to maintain or transpose access permissions to ensure only authorized individuals have access to specific datasets, files or buckets.

With metadata being pervasive nowadays, especially with object storage, metadata attributes must also be maintained and data integrity must be carried over from source to the target system.

System utilities are not able to maintain or transpose those attributes between source and target systems. They also remain limited from a manageability and scalability perspective, not to mention incapable of performing cross-protocol migration, making them ill-suited for enterprise-grade migration projects involving critical data and systems.

How Interlock HelpsEnter Interlock, a private company founded in 2009 with the precise mission of accelerating enterprise-grade, complex, cross-protocol, and at scale data migration with stringent compliance requirements.

Interlock’s expertise in the data migration space is consumable via two different service offerings:

  • A3 DATAFORGE Classic which is available as a fully managed professional services engagement offering the most comprehensive level of support, even for the most complex projects. This includes use of commercial and custom-built data connectors to ensure seamless migration while maintaining the data consonance imperatives highlighted earlier.
  • A3 DATAFORGE is a software data migration solution usable by certified professionals. Compared to DATAFORGE Classic, this solution offers a self-service turnkey approach that is suitable for most migration activities, with the ability to consume professional services as an addition to the licensing costs.

From an architectural standpoint, the solution consists of two pieces – a DATAFORGE Application Node, which acts as the management plane, and data movers. Those components are deployed as virtual machine templates (OVA files) on VMware vSphere.

ScreenshotWith A3 DATAFORGE and A3 DATAFORGE Classic, migrations are executed in five distinct phases:

  • Planning phase: This initial phase entails mapping of application and data types, understanding the impact of migration activities, defining the data migration strategy, and identifying timelines and any potential challenges.
  • Pre-migration phase: A preparatory phase that is all about assessment and risk mitigation. Some activities include mapping of initial source/target, backups and dry run tests, resolution of potential issues and implementation of secure data migration requirements.
  • Migration execution phase: Data migration activities occur with data movers being actively used to parallelize data transfers and achieve higher throughput rates compared to the standard copy tools. If necessary, gateways are bypassed, and in-flight optimizations are made to maintain consistency during cross-protocol migrations.
  • Post-migration phase: Post-migration, data validation and verification activities are performed to ensure that there was no loss or corruption. Additionally, application functionality is tested in this phase, and, in the case of subsequent migration activities, performance too is measured and tuned.
  • Completion and Reporting phase: This last phase involves final health and consistency checks, delivery of migration reports, and handover activities.

Key Differentiators of A3 DATAFORGEA3 DATAFORGE is one of the many data migration solutions available to organizations. But there are some big differentiators that sets it apart from competing solutions.

  • Cross-protocol Migration and Application Consonance – This ensures that applications remain functional even during migration, and that all data and metadata attributes remain unaltered to avoid any application-level changes.
  • Scalability and Throughput – The solution is designed with massive scalability and throughput in mind. Data movers can operate with up to 128 threads per mover, and they can be deployed in as many numbers as needed to parallelize operations, making bandwidth between source and target the only limiting factor.
  • Integration and Connectors – In addition to providing standard connectors for the most popular applications, the company also offers custom-built connectors allowing migration activities to reach a much deeper level of data awareness, thus delivering much better data consistency.
  • Storage Layer Migrations and Data Gateway Bypass – This significantly speeds up data migration activities compared to native data migration tools or system utilities
  • Legacy System Expertise: Although perfectly capable of managing modern workloads, Interlock Technology has considerable experience with migration of legacy systems, allowing it to transform data while keeping compliance attributes intact.

ConclusionIntegrity and consistency remain essential requirements for any dataset, structured or unstructured. Yet they are often overlooked in migration activities simply because it is a given.

With cross-protocol migrations (often happening when legacy source systems are being retired and replaced with modern incompatible solutions), there is no simple or straightforward way to ensure that data consistency and ancillary data attributes, permissions, and integrity stamps are maintained.

For such migrations, the capabilities delivered by Interlock’s A3 DATAFORGE solutions provide not only assurance that data migration projects will be a success, but also surety that data consistency and consonance are maintained throughout the application’s lifecycle.

For more, head over to the Tech Field Day website or YouTube channel to see videos from the Interlock Tech Field Day Showcase.


© Gestalt IT, LLC for Gestalt IT: Maintaining Data Consistency in Migration Projects with Interlock

View Details

Far from their history in data backup, Commvault is firmly focused on enabling modern, cloud-first businesses to achieve continuous operation. This Gestalt IT Tech Talks roundtable discussion follows Commvault’s October 2024 Shift event in London, with attendees Stephen Foskett, Brian Booden, Jack Poller, and Tom Hollingsworth. Commvault has a decades-long history in protecting data for archive and backup, and is one of the most successful companies in this business, especially for larger enterprises. But the company has radically shifted its messaging and product development to enable business applications to recover from cyber attacks and other incidents. Commvault is also very focused on applications outside the datacenter, notably cloud-native applications, SaaS platforms, and even edge environments. This is reflected by recent acquisitions of Appranix and Clumio, which were featured at the event, and development of the rest of the Commvault stack on the foundation of their Metallic initiative. The attendees agreed that Commvault truly is shifting its approach and products to enable continuous business.


Commvault Angles towards Cloud-Focused Cyber ResilienceEvery calendar year, cybersecurity vendors compete for the spotlight in the cyber resilience category. The total addressable market (TAM) which is projected to cross $300 billion in the next three years according to reports, sees a torrent of new security-assured point products land every few months. But companies still face an unusual degree of risk from threat actors.

A majority of these providers focus on prevention as the central theme of ransomware protection. At Commvault SHIFT London 2024, Commvault called this out as misplaced focus and proposed a new way of addressing the issue. CEO, Sanjay Mirchandani phrased it as “continuous business”.

Business interruptions attributed to cyberattacks wreak havoc across sectors globally. These events are a strong reminder of business’ rising cyber debt and the cascading effect it can have when an external actor exploits it to seize control of the data or the infrastructure.

Ransomware attacks and cyber exploits have resulted in business operation crashes, leading to halting of service availability worldwide. Almost always a fix is dispatched and with some luck, things return to normal but not before a certain time window during which everything hangs in the limbo. That period, according to Statista, is 24 days on average in case of a ransomware incident. Every minute of this downtime costs real dollars.

“The cyber security industry as a whole is focused on defense and preventing attacks than on what happens after an attack occurs,” said Jack Poller, analyst, echoing Commvault’s point while reflecting on the highlights of SHIFT that he attended virtually, during a Gestalt IT Tech Talk Roundtable discussion.

Continuous business refers to an always-on state where business is operational and available around the clock even after a disruptive event. The idea is premised on four principles – continuous security, continuous readiness, continuous recovery and continuous rebalance – a feat, if accomplished, will lead to continued operations through a crisis.

“Disaster recovery is no longer about disaster recovery from a mechanical or equipment failure or some sort of IT related issue. It’s from anything that takes your digital infrastructure down and that change in attitude is critical to thinking that we need to keep the business operating and we can’t separate cybersecurity attacks from any other IT issue that causes failure. Commvault’s taking the position that businesses need to continue operating regardless of what the cause of a catastrophe is, reflects that,” Poller commented.

Outmoded and enormously costly architectural constructs in the cloud-first century is the biggest impediment of businesses trying to get to a cyber-ready state.

In an open letter to companies, Mirchandani wrote, “This lack of vision and ambition comes at the expense of the customer and the vulnerability of the business. It’s a decision to stay the profitable, legacy course because of the belief that the effort outweighs the reward.”

Commvault, under the leadership of Mirchandani and a newly formed management team, is angling to have the old, overused playbook replaced with a new cloud-focused strategy. At SHIFT London, the company declared a decisive changeover from backup-and-recovery to setting a new course in data protection and infrastructure resilience.

Enabling Cyber Readiness and Resilience in a Cloud-First WorldMany of Commvault’s prominent releases from the event rely on this new philosophy, including Cloud Rewind which was the crown jewel of the event.

Added to the Commvault Cloud Platform, Cloud Rewind is a solution that borrows advanced recovery and rebuild capabilities of Appranix which Commvault snapped up in a recent acquisition, and delivers quick automated cyber recovery of entire cloud applications and data environments. Commvault calls it a “cloud time machine” that can make organizations return to business as usual in just minutes after an attack.

The transition to hybrid multi-cloud has long been predicted, but businesses are encountering significant headwinds with protecting and managing data and applications across widely dissimilar platforms.

“Businesses don’t just sit in Google or AWS or one of these systems,” said Brian Booden, panelist, and CEO of DataGlow IT. “They’re spreadeagled across all of them and one of the quotes that resonated with me is that “if your business is in the cloud, cloud is your business”. That’s just extremely relevant today now as companies start to move off-prem and into the cloud.”

Commvault’s broad support for the hybrid multi-cloud includes GCP support, and now offerings on AWS. A vast number of AWS customers leverage Commvault’s data protection capabilities on the platform. Seeing the number grow continually, Commvault announced the complete Commvault Cloud Platform for AWS. This will allow customers on AWS to leverage Cloud Rewind for S3 data, Air Gap Protect for immutability, and Commvault’s Cleanroom Recovery solution for test and recovery of data assets.

Commvault further strengthened its cloud-first initiative with the acquisition of Clumio, a cloud-native data protection company that lends it easy-to-consume and powerful backup and recovery capabilities, besides a spectrum of security features like air-gapped backup and encryption.

The third biggest release was a joint compliance solution with Pure Storage. In partnership with Pure Storage, Commvault has developed an integrated cyber solution aimed at assisting customers in the finance sector to meet the compliance requirements and regulations of Digital Operational Resilience Act (DORA).

Wrapping UpOverall, the theme of cloud-focused cyber resilience and continuous business resonated throughout the event, and it will likely be the North Star for Commvault’s journey here on out. As Poller rightly said, “Commvault is saying wherever your infrastructure lives and wherever you have critical data, we’re going to protect it. That message came across loud and clear.”

Check out the Tech Talk Roundtable to hear the panel’s impression of the announcements from SHIFT London 2024. For more conversations like it, catch Stephen Foskett in conversation with Max Mortillaro and Chris Evans at SHIFT, and Daniel Newman and Patrick Moorehead’s sit-down with Mirchandani. Keep your eyes peeled for more articles on Commvault SHIFT publishing at Gestalt IT.


Tech Talk Information:Stephen Foskett is the Organizer of the Tech Field Day Event Series, now part of The Futurum Group. Connect with Stephen on LinkedIn or on X/Twitter.

Tom Hollingsworth is the Networking Analyst for The Futurum Group and Event Lead for Tech Field Day. You can connect with Tom on LinkedIn and X/Twitter. Find out more on his blog or on the Tech Field Day website.

Jack Poller is and industry leading cybersecurity analyst and Founder of Paradigm Technica. You can connect with Jack on LinkedIn or on X/Twitter. Learn more on Paradigm Technica’s website.

Brian Booden is the Managing Director at DataGlow IT. You can connect with Brian on X/Twitter and on LinkedIn. Learn more about DataGlow IT on their website.


Gestalt IT and Tech Field Day are now part of The Futurum Group.


© Gestalt IT, LLC for Gestalt IT: Commvault Shifts to Focus on Continuous Business

View Details

Ransomware is a critical challenge to business security today. I don’t need to tell you how horrifying it is when you find out that your systems have been infected with a strain of malware designed to knock you offline or make you pay some exorbitant extortion fee to free your data from being encrypted. Why are these attacks so effective? Because businesses can’t be offline. When your business isn’t selling to customers you are losing money. The malware writers know that. They attack in the hopes that paying off the gang is less money than you will lose if you can’t satisfy your customer base.

Recently, I was able to be a part of Commvault Shift as a remote live blogger. I watched as the company outlined the way they handle data protection and business continuity in the modern world. I found it interested that the three concepts they outlined for continuous business, namely Cost, Complexity, and Control, could also apply to reasoning behind why attackers succeed in taking a business offline:

  • Cost: Does anyone know what it costs today to bring a business back from a ransomware infection? I’m not talking about the downtime either. How about the hours spent restoring data and provisioning services? The idea of shutting everything off and trying to turn it back on again gives many organizations cold sweats.
  • Complexity: Don’t forget the extra troubleshooting that has to happen when your old technical debt doesn’t play well with your new projects. You may have built your organization on the latest cloud-first platforms but there’s no guarantee that some old software program isn’t going to cause a problem when it fails to connect to an obsoleted API. All of that means more time for your people to sort out the glitches, which adds to your outage.
  • Control: Who is really in charge here? You? Or the attackers? If they have inroads into your organization you could find yourself facing wave after wave of attacks even after a successful restoration. Something as simple as creating a backdoor account in Active Directory before knocking it offline could mask a foothold that gives them access to wreak havoc in the future.

Thankfully, Commvault knows how hard these issues are to solve in the world of restoration. Even when you’re not fighting against an adversarial group you have to figure out how to make the best use of your resources before getting everyone back to a known-good state. You need expertise. You need technology designed to meet your modern needs. And you need it to work even if your people are busy doing other things.

Go Back to Go ForwardThat’s why Cloud Rewind was so exciting to me. Formerly known as Appranix before being acquired back in April 2024, Cloud Rewind does all the things that backup and restore vendors have been promising for years. Yes, it does back up the data in your cloud environment. But it also maps out cloud services dependencies and looks for configurations that have drifted away from your standards. Because it is aware of the platform configuration it will not only restore the data but the state of the system when it was backed up. No more guessing about how the system was configured in the first place. Cloud Rewind just restores it all.

You can configure Cloud Rewind to take backups of the system as often as every five minutes. And with enough storage you can create a repository that allows you to go back as far as possible within the limits of physics. These frequent backups mean that you can find the perfect balance between the need to retain data and the possible dwell time of your attackers. The scariest part of restoring the data isn’t when you find out you’ve been breached. It’s figuring out how long they were in the system before you found them, or before they attacked you. Combined with the other great security features in the Commvault suite you can be assured that the data being restored by Cloud Reward is secure and uninfected.

The best part? You can download Cloud Reward directly from the AWS Marketplace today and start using it. Use costs are per instance per hour but the peace of mind that you get from knowing your data isn’t going to be encrypted or sold to someone else is well worth the investment. The cloud isn’t some magical platform that keeps everything forever. You need to have a solution that can get you back in business. Don’t count on old technology doing that in a world of cloud complexity. Use a tool like Commvault Cloud Rewind and leave your worries for the other hard stuff.


© Gestalt IT, LLC for Gestalt IT: The Security of Availability from Commvault Shift

View Details

Commvault Shift 2024 is exploring how Commvault is redefining the way we think about approaching cyber resilience for the AI era. We invite you to follow our live blog of the event or register here.


And that's a wrap. Your 5 key takeaways from #CommvaultShift pic.twitter.com/wSWESlaXqj

— Karen Lopez (@datachick) October 9, 2024

Ethan Banks on LinkedIn: https://www.linkedin.com/posts/ethanbanks_commvaultshift-activity-7249878364717494273-XAob/?utm_source=share&utm_medium=member_desktop

Key Takeaways#CommvaultSHIFT #ContinuousBusiness #CyberResilience @Commvault pic.twitter.com/n1dvInYlOf

— Jack L. Poller (@poller) October 9, 2024

6 Principles of Resilient Systems. #CommvaultSHIFT pic.twitter.com/6WJkynBz5s

— Karen Lopez (@datachick) October 9, 2024

What makes it a competitive advantage: complex systems are hard to duplicate.#CommvaultSHIFT #ContinuousBusiness #CyberResilience @Commvault

— Jack L. Poller (@poller) October 9, 2024

Justin Warren on Mastodon: https://eigenmagic.net/@daedalus/113279280236576925

Jack Poller on Mastodon: https://infosec.exchange/@poller/113279310249245425

Reeves is using the human immune system as an example of a resilient system@Commvault#CommvaultSHIFT #ContinuousBusiness #CyberResilience

— Jack L. Poller (@poller) October 9, 2024

Resilience is formed of 6 principle highlights @MartinKReeves @TechFieldDay @Commvault #CommvaultShift pic.twitter.com/CFMuc9JXbd

— Brian Booden (@brian_booden) October 9, 2024

Justin Warren on Mastodon: https://eigenmagic.net/@daedalus/113279272462638481

Embeddedness: making sure your resilient system is embedded in a resilient system@Commvault#CommvaultSHIFT #ContinuousBusiness #CyberResilience

— Jack L. Poller (@poller) October 9, 2024

The 6 principles of resilient systems@Commvault#CommvaultSHIFT #ContinuousBusiness #CyberResilience pic.twitter.com/ccS4aQVztq

— Jack L. Poller (@poller) October 9, 2024

4th phase: instead of recovering to the current state, reimagine possibilities to recover to a new, better state@Commvault#CommvaultSHIFT #ContinuousBusiness #CyberResilience

— Jack L. Poller (@poller) October 9, 2024

Interesting take on the 4 phases of resiliency@Commvault#CommvaultSHIFT #ContinuousBusiness #CyberResilience pic.twitter.com/juRtw5xq90

— Jack L. Poller (@poller) October 9, 2024

Jack Poller on Mastodon: https://infosec.exchange/@poller/113279274158247487

Final session is Martin Reeves author of "The Imagination Machine" talking about resilience as a competitive advantage@Commvault #CommvaultSHIFT #ContinuousBusiness #CyberResilience pic.twitter.com/QmoJqeyJVH

— Jack L. Poller (@poller) October 9, 2024

Panel is talking about typical data security concerns re: data classification and how that might be used by AI.

What about the massive volume of data generated by AI that is unreproducible?@Commvault #CommvaultSHIFT #ContinuousBusiness #CyberResilience

— Jack L. Poller (@poller) October 9, 2024

More at the link below!Ethan Banks on LinkedIn: https://www.linkedin.com/posts/ethanbanks_commvaultshift-activity-7249861825406607361-DX1t/?utm_source=share&utm_medium=member_desktop

Bonus session is another panel discussing AI.

First up, how are people using AI: for offensive and defensive activities@Commvault #CommvaultSHIFT #ContinuousBusiness #CyberResilience

— Jack L. Poller (@poller) October 9, 2024

Getting ready for a bonus session on AI and Cyber Resilience@Commvault #CommvaultSHIFT #ContinuousBusiness #CyberResilience

— Jack L. Poller (@poller) October 9, 2024

"I work in the fastest and slowest moving industry in the world" We cannot even change passwords on one side, and on the other we have the potential of AI and Quantum Computing fuelled. Resilience is very important! @TechFieldDay @Commvault #CommvaultShift pic.twitter.com/0UvFOjZ20V

— Brian Booden (@brian_booden) October 9, 2024

Stephen Foskett on LinkedIn: https://www.linkedin.com/posts/sfoskett_commvaultshift-activity-7249841116894695424-43Xw/?utm_source=share&utm_medium=member_ios

Justin Warren on Mastodon: https://eigenmagic.net/@daedalus/113279064402837161

AI is a key business imperative, but letting it run along unregulated isn't a good idea either. #CommvaultSHIFT

— Tom Hollingsworth (@NetworkingNerd) October 9, 2024

Panel: Now talking about regulating AI – governments see regulation as a way for nation/states to gain control and economic advantage@Commvault #CommvaultSHIFT #ContinuousBusiness #CyberResilience

— Jack L. Poller (@poller) October 9, 2024

Justin Warren on Mastodon: https://eigenmagic.net/@daedalus/113279021182402976

Criminal do have access to the same tools that we do. The upside? We know how they will use them. #CommvaultSHIFT

— Tom Hollingsworth (@NetworkingNerd) October 9, 2024

I'm curious about DORA. It keeps coming up ominously. #CommvaultSHIFT

— Tom Hollingsworth (@NetworkingNerd) October 9, 2024

"Handling DORA isn't a project, it's a transformation" @darrencthomson with a truth bomb about how seriously we need to take new AI and Data governance regulations. @TechFieldDay @Commvault #CommvaultShift pic.twitter.com/x3IEMrJd31

— Brian Booden (@brian_booden) October 9, 2024

Panel: 1st step is make a list of your crown jewels.

Reinforces my thesis that the #1 cybersecurity problem is visibility: you can't protect what you don't know about@Commvault #CommvaultSHIFT #ContinuousBusiness #CyberResilience

— Jack L. Poller (@poller) October 9, 2024

Jack Poller on Mastodon: https://infosec.exchange/@poller/113278992860392085

Justin Warren on Mastodon: https://eigenmagic.net/@daedalus/113278976335127041

Panel discussion on Compliance with Harvard Business Review #CommvaultSHIFT #ContinuousBusiness #CyberResilience @Commvault pic.twitter.com/fn93D8AKoM

— Jack L. Poller (@poller) October 9, 2024

Justin Warren on Mastodon: https://eigenmagic.net/@daedalus/113278928043332026

Article Link: https://pivotnine.com/the-crux/archive/swearing-in-source-code/

An impressive suite of announcements summarised by @ksrajiv but headlined by the Craig David ?endorsed Cloud Rewind, enhanced S3 and @awscloud support, AD Forest Recovery (Wow!) and @Google Workspace support. @TechFieldDay @Commvault #CommvaultShift pic.twitter.com/uIwEFzK159

— Brian Booden (@brian_booden) October 9, 2024

"We just want stuff to work, we don't want to think about it. If it doesn't, patients will ultimately suffer" @UmangPatel1 Chief Clinical Information Officer at @Microsoft lays out the importance of a speedy and accurate recovery. @TechFieldDay @Commvault #CommvaultShift pic.twitter.com/CZDzAW2olV

— Brian Booden (@brian_booden) October 9, 2024

Justin Warren on Mastodon: https://eigenmagic.net/@daedalus/113278913822437917

Justin Warren on Mastodon: https://eigenmagic.net/@daedalus/113278788833474012

Justin Warren on Mastodon: https://eigenmagic.net/@daedalus/113278783659368946

Justin Warren on Mastodon: https://eigenmagic.net/@daedalus/113278740102740573

Justin Warren on Mastondon: https://eigenmagic.net/@daedalus/113278713802559266

News https://buff.ly/3XTIvo2Commvault announces support for #AWS for Air Gap protect, Cleanroom Recovery.#CommvaultSHIFT

— Karen (Kitty) Lopez (@datachick.bsky.social) 2024-10-09T18:24:50.712Z

Ethan Banks on LinkedIn: https://www.linkedin.com/posts/ethanbanks_commvaultshift-activity-7249846749987536896-HCs_/?utm_source=share&utm_medium=member_desktop

The prospect of Air Gap Protect and Cleanroom Recovery being native @awscloud EC2 instances is quite interesting. @TechFieldDay @Commvault #CommvaultShift pic.twitter.com/dUHH77FzlN

— Brian Booden (@brian_booden) October 9, 2024

Connecting with new faces and reconnecting with my @TechFieldDay OG peeps @Commvault SHIFT in London! Always fun when @SFoskett is around. #CommvaultSHIFT #techfielday @ArchitectingIT @maxmortillaro ?? pic.twitter.com/pkCkHSeIS0

— Colleen wants to see The Hives in concert!? (@colleencoll) October 9, 2024

Commvault integrates w/ @PaloAltoNtwks XSOAR to automatically spin up a clean room during an incident to accelerate recovery#CommvaultSHIFT #ContinuousBusiness #CyberResilience @Commvault

— Jack L. Poller (@poller) October 9, 2024

Major enhancements to Air Gap Protect and Cleanroom Recovery on @awscloud @TechFieldDay @Commvault #CommvaultShift pic.twitter.com/sC1cbBmf84

— Brian Booden (@brian_booden) October 9, 2024

Factory reset enables recovery of host to clean copy w/ no risk of malware/ransomware

Proactive and automated resiliency solution#CommvaultSHIFT #ContinuousBusiness #CyberResilience @Commvault

— Jack L. Poller (@poller) October 9, 2024

Cleanroom now supports AWS – data is closer and you can specify AWS as a cleanroom target to bring up your initial environment during recovery#CommvaultSHIFT #ContinuousBusiness #CyberResilience @Commvault

— Jack L. Poller (@poller) October 9, 2024

50% of organizations with remote operations have experience a cyberattack. Increased footprint means more entry points. #CommvaultSHIFT

— Tom Hollingsworth (@NetworkingNerd) October 9, 2024

93% of businesses don't have a reliable backup

#CommvaultSHIFT #ContinuousBusiness #CyberResilience @Commvault

— Jack L. Poller (@poller) October 9, 2024

Announcing HyperScale X Compute and Edge: a scale-out solution that combines data protection software, compute, operating system, and storage to simplify data protection and backup management @TechFieldDay @Commvault #CommvaultShift pic.twitter.com/IQbCfpUNui

— Brian Booden (@brian_booden) October 9, 2024

HyperScale Edge is hyperconvered, but HyperScale Compute doesn't include storage. #CommvaultSHIFT #ContinuousBusiness #CyberResilience @Commvault

— Jack L. Poller (@poller) October 9, 2024

HyperScale X, HyperScale Computer, HyperScale Edge – various hyperconvered(?) solutions for cyber resiliency that can be centrally managed#CommvaultSHIFT #ContinuousBusiness #CyberResilience@Commvault

— Jack L. Poller (@poller) October 9, 2024

Jack Poller on Mastodon: https://infosec.exchange/@poller/113278799940326539

Justin Warren on Mastodon: https://eigenmagic.net/@daedalus/113278376406634479

Jack Poller on Mastodon: https://infosec.exchange/@poller/113278799940326539

Ethan Banks on LinkedIn: https://www.linkedin.com/posts/ethanbanks_commvaultshift-activity-7249842272639299584-8VQ4/?utm_source=share&utm_medium=member_desktop

Pranay Ahlawat explains the breadth and depth of support for @Microsoft 365, and the new abilities across @Google Workspace @TechFieldDay @Commvault #CommvaultShift pic.twitter.com/K2GElzr6E6

— Brian Booden (@brian_booden) October 9, 2024

Commvault announces the ability to recover AD forests — this is really hard and complicated, takes days when doing it manually. Only 2 other vendors can do this!@Commvault #CommvaultSHIFT #ContinuousBusiness #CyberResilience pic.twitter.com/3t3egfDmSa

— Jack L. Poller (@poller) October 9, 2024

Commvault claims that 9/10 attacks are targeted at Active Directory@Commvault #CommvaultSHIFT #ContinuousBusiness #CyberResilience

— Jack L. Poller (@poller) October 9, 2024

Don't assume your SaaS vendor is backing up your data. Sure, it may not be on a dead drive but how often does it get accidentally deleted? #CommvaultSHIFT

— Tom Hollingsworth (@NetworkingNerd) October 9, 2024

Commvault now protects Google Workspace in addition to MSFT M365@Commvault #CommvaultSHIFT #ContinuousBusiness #CyberResilience pic.twitter.com/HgvCoc01AF

— Jack L. Poller (@poller) October 9, 2024

"The cost is a fraction of using the large hyperscalers" Venkata Sudhakar Nagandla with a bold statement @TechFieldDay @Commvault #CommvaultShift pic.twitter.com/caaXbvI3fZ

— Brian Booden (@brian_booden) October 9, 2024

Commvault announces new support for
Azure Data Lake Gen 2
Databricks
Cosmos DB
Mongo DB
RAG extensions for Postgres
AWS S3

via Cloud Rewind#cloud pic.twitter.com/s3UbTeEFW5

— Karen Lopez (@datachick) October 9, 2024

Stephen Foskett on LinkedIn: https://www.linkedin.com/posts/sfoskett_commvaultshift-activity-7249841116894695424-43Xw/?utm_source=share&utm_medium=member_desktop

Cloud Rewind is the evolution of the Appranix acquisition and helps you deal w/ config drift, cyber-attacks, and infrastructure failures@Commvault #CommvaultSHIFT #ContinuousBusiness #CyberResilience

— Jack L. Poller (@poller) October 9, 2024

Cloud Rewind is more than data recovery – it's the ability to recover the infrastructure: apps, storage, network, and data@Commvault #CommvaultSHIFT #ContinuousBusiness #CyberResilience

— Jack L. Poller (@poller) October 9, 2024

Platform https://t.co/dsYEsTrReg

Moving from just backups to "Continuous business" by embracing multi-cloud, working with Azure, AWS, and Google Clouds.

Today, they are announcing support for AWS for Commvault Cloud. pic.twitter.com/kWNuTFgDPt

— Karen Lopez (@datachick) October 9, 2024

Over 70% of cloud resources are not protected¹. This means that recovering and rebuilding all cloud configurations after an attack takes an average of 24 days and costs organizations billions in lost revenue and brand reputation. https://t.co/OOK52pkdXs pic.twitter.com/3YvKpaAskd

— Karen Lopez (@datachick) October 9, 2024

Ethan Banks on LinkedIn: https://www.linkedin.com/posts/ethanbanks_commvaultshift-activity-7249839020585721856-cgAP?utm_source=share&utm_medium=member_desktop

Jack Poller on Mastodon: https://infosec.exchange/@poller/113278717856816033

Cloud Rewind…when the Cloud says "yo, rewind me" ?@TechFieldDay @Commvault #CommvaultShift #CraigDavidLivesOn pic.twitter.com/K8UGylORI0

— Brian Booden (@brian_booden) October 9, 2024

Justin Warren on Mastodon: https://eigenmagic.net/@daedalus/113278376406634479

Commvault claims they now recover at 2.5TB/hr, 6x improvement in the last year@Commvault #CommvaultSHIFT #ContinuousBusiness #CyberResilience

— Jack L. Poller (@poller) October 9, 2024

Over 3 weeks to recover from an outage? Could you go 3 weeks without getting paid? #CommvaultSHIFT

— Tom Hollingsworth (@NetworkingNerd) October 9, 2024

Here are a few more photos from #CommvaultShift live in London yesterday! @Commvault @ClumioInc pic.twitter.com/MNYR3fUufO

— Stephen Foskett (@SFoskett) October 9, 2024

Let's not underestimate the impact of TCO when discussing these solutions. Let's see how much that is referenced moving forwards in the announcements @TechFieldDay @Commvault #CommvaultShift pic.twitter.com/cpweOrj3jy

— Brian Booden (@brian_booden) October 9, 2024

5 key challenges for protecting data in the cloud

full stack environment, TCO, Security/Compliance, fragmentation, recovery of infra + data@Commvault #CommvaultSHIFT #ContinuousBusiness #CyberResilience pic.twitter.com/5DnOrP49rw

— Jack L. Poller (@poller) October 9, 2024

Now it's @ksrajiv "24 days to recover from a compromised system" That's a loooong time for large orgs. @TechFieldDay @Commvault #CommvaultShift pic.twitter.com/yh3Xol5TZP

— Brian Booden (@brian_booden) October 9, 2024

Joep Piscaer on LinkedIn: https://www.linkedin.com/posts/jpiscaer_commvaultshift-activity-7249804692652707842-tnUE/?utm_source=share&utm_medium=member_ios

Stephen Foskett on LinkedIn: https://www.linkedin.com/posts/sfoskett_commvaultshift-activity-7249836139342106625-6rfP/?utm_source=share&utm_medium=member_ios

Ethan Banks on LinkedIn: https://www.linkedin.com/posts/ethanbanks_commvaultshift-activity-7249833457835499543-cgf0?utm_source=share&utm_medium=member_desktop

Justin Warren on Mastodon: https://eigenmagic.net/@daedalus/113278376406634479

Jack Poller on Mastodon: https://infosec.exchange/@poller/113278575155991656

Jumping to the next #CommvaultSHIFT session: Innovating Our Platform To Create a New Paradigm

— Tom Hollingsworth (@NetworkingNerd) October 9, 2024

In summary, @ClumioInc is bringing @awscloud S3 Bucket Version Control to @Commvault to easily be able to restate your bucket position post-corruption @TechFieldDay @Commvault #CommvaultShift pic.twitter.com/dLDj9ueatk

— Brian Booden (@brian_booden) October 9, 2024

Justin Warren on Mastodon: https://eigenmagic.net/@daedalus/113278376406634479

SHIFT Virtual | 2024 https://t.co/4dEjvU05HE

Still time to join us for announcements from Commvautl pic.twitter.com/eZ5FRgbUXL

— Karen Lopez (@datachick) October 9, 2024

The Clumio solution is helping Commvault expand from Azure into AWS (and hopefully, GCP)@Commvault #CommvaultSHIFT #ContinuousBusiness #CyberResilience

— Jack L. Poller (@poller) October 9, 2024

With Clumio, you can manage S3 object versioning stack, rolling back to any point in time, for any scale of S3 – another form of Cloud Rewind@Commvault #CommvaultSHIFT #ContinuousBusiness #CyberResilience

— Jack L. Poller (@poller) October 9, 2024

Ethan Banks on LinkedIn: https://www.linkedin.com/posts/ethanbanks_commvaultshift-activity-7249832474522857472-D72g/?utm_source=share&utm_medium=member_desktop

Justin Warren on Mastodon: https://eigenmagic.net/@daedalus/113278376406634479

Love hearing from the @ClumioInc folks during #CommvaultSHIFT!

— Tom Hollingsworth (@NetworkingNerd) October 9, 2024

Appranix is now Cloud Rewind. I like it. #CommvaultSHIFT

— Tom Hollingsworth (@NetworkingNerd) October 9, 2024

Cloud Rewind is a statement announcement, and now we get to meet @ClumioInc, who acquired to help solve cyber recovery challenges on the public cloud. @TechFieldDay @Commvault #CommvaultShift pic.twitter.com/aq47SRqxDn

— Brian Booden (@brian_booden) October 9, 2024

Ethan Banks on LinkedIn: https://www.linkedin.com/feed/update/urn:li:activity:7249830637753245696/

Cloud Rewind (leveraging Appranix acquisition) enables restoration of your infrastructure and data – go back to last known good state@Commvault #CommvaultSHIFT #ContinuousBusiness #CyberResilience pic.twitter.com/WDu9ODVEXC

— Jack L. Poller (@poller) October 9, 2024

Commvault now has a new built-in cyber resilience dashboard to analyze your readiness and recovery capabilities@Commvault #CommvaultSHIFT #ContinuousBusiness #CyberResilience

— Jack L. Poller (@poller) October 9, 2024

Continuous security is something you should be doing all the time. This isn't a checkbox, it should be a way of life. #CommvaultSHIFT

— Tom Hollingsworth (@NetworkingNerd) October 9, 2024

"Cyber recovery should not be a guessing game" as we discuss Cleanroom Recovery @TechFieldDay @Commvault #CommvaultShift pic.twitter.com/GvKjRj8l4e

— Brian Booden (@brian_booden) October 9, 2024

Continuous business is a big driver for today's resilience initiatives. Remember when ESPN and HBO didn't do programming after 2am? #CommvaultSHIFT

— Tom Hollingsworth (@NetworkingNerd) October 9, 2024

Over 1,000 people tuning in on the #CommvaultSHIFT live stream. Quite impressive!

— Tom Hollingsworth (@NetworkingNerd) October 9, 2024

Fewer than half of all companies are confident in their recovery plans; 20% don't test recovery at all@Commvault #CommvaultSHIFT #ContinuousBusiness #CyberResilience

— Jack L. Poller (@poller) October 9, 2024

Continous Rebalancing leverages mirroring your data to all your clouds so you can make necessary adjustments@Commvault #CommvaultSHIFT #ContinuousBusiness #CyberResilience pic.twitter.com/DNXNuVkvHi

— Jack L. Poller (@poller) October 9, 2024

Jack Poller on Mastodon: https://infosec.exchange/@poller/113278575155991656

Justin Warren on Mastodon: https://eigenmagic.net/@daedalus/113278376406634479

Say Goodbye to old assumptions, say Hello to Commvault Cloud – it's your cloud, Commvault is there to help@Commvault #CommvaultSHIFT #ContinuousBusiness #CyberResilience pic.twitter.com/xN2lKzeig0

— Jack L. Poller (@poller) October 9, 2024

"One size does NOT fit all" Gen AI datasets needs a different approach to Cloud Native and SaaS applications @TechFieldDay @Commvault #CommvaultShift pic.twitter.com/e9y2EWFJxe

— Brian Booden (@brian_booden) October 9, 2024

Jack Poller on Mastodon: https://infosec.exchange/@poller/113278566885370245

"Own Your Cloud – If your Business is in the cloud, then CLOUD is your business" Straight shooting from @mirchi111 but I like It @TechFieldDay @Commvault#CommvaultShift pic.twitter.com/UgI3Tmp8oF

— Brian Booden (@brian_booden) October 9, 2024

Commvault is shifting to the concept of Continuous Business – always on availability and resilience@Commvault #CommvaultSHIFT #ContinuousBusiness #CyberResilience pic.twitter.com/cvXd13hjIc

— Jack L. Poller (@poller) October 9, 2024

First up is Sanjay Mirchandani! He always cuts an impressive figure in these keynotes. #CommvaultSHIFT

— Tom Hollingsworth (@NetworkingNerd) October 9, 2024

The first session is Reimagining Resilience for the Cloud-First Enterprise. Can I just say I love "resilience"? #CommvaultSHIFT

— Tom Hollingsworth (@NetworkingNerd) October 9, 2024

"Bolted on processes on existing data process" @mirchi111 describes the state of play before we enter the era of Continuous Business @TechFieldDay @Commvault #CommvaultShift pic.twitter.com/dOPjWveAQ7

— Brian Booden (@brian_booden) October 9, 2024

Commvault SHIFT is starting!@Commvault #CommvaultSHIFT #ContinuousBusiness #CyberResilience

— Jack L. Poller (@poller) October 9, 2024

In a few minutes, it's time to join @Commvault #CommvaultSHIFT, the premier cyber resiliency event that offers a radical shift in perspective, helping orgs embrace cyber readiness and recoverability for always-on continuous business. I'm on deck representing @TechFieldDay? pic.twitter.com/QrxZKl18EM

— Brian Booden (@brian_booden) October 9, 2024

I'm going to be live blogging the #CommvaultShift video stream here today starting at 1pm ET! You can also keep up with all the other coverage by checking out this page: https://t.co/AHryx6ZMhX

— Tom Hollingsworth (@NetworkingNerd) October 9, 2024

Only 1 hour from the start of Commvault SHIFT, where Commvault will describe their shift in perspective to cyber resiliency in a cloud-first world#CommvaultSHIFT @Commvault #ContinuousBusiness #CyberResilience

— Jack L. Poller (@poller) October 9, 2024

I attended #CommvaultSHIFT yesterday, and I'll be covering their cyber resilience and business continuity livestream in this thread. I'm hoping to hear more about the recent @Clumio acquisition, as well as integrations of @Appranix into the portfolio. Let's dive in! pic.twitter.com/G4wZ9JDj3u

— Joep Piscaer (@jpiscaer) October 9, 2024

Colleen Coll: https://www.linkedin.com/posts/colleen-coll-b971505_dataprotection-commvaultshift-cyberresilience-activity-7249801934168084480-yVWx/?utm_source=share&utm_medium=member_desktop

Read more on Six Five Media: https://sixfivemedia.com/all-videos/shifting-to-cloud-first-cyber-resilience-with-commvault/

ScreenshotStephen Foskett on LinkedIn: https://www.linkedin.com/pulse/commvault-embraces-shift-continuous-business-2024-stephen-foskett-ssuae/

Gestalt IT on LinkedIn: https://www.linkedin.com/posts/gestalt-it_cybersecurity-cyberresilience-commvaultshift-activity-7249414742950178819-KOf1/?utm_source=share&utm_medium=member_ios

Gestalt IT on LinkedIn: https://www.linkedin.com/posts/gestalt-it_commvaultshift-activity-7249060031449436160-Vf36/?utm_source=share&utm_medium=member_ios

Jack Poller on LinkedIn: https://www.linkedin.com/posts/jackpoller_commvaultshift-continuousbusiness-cyberreadiness-activity-7249465839488221185-Kp9I?utm_source=share&utm_medium=member_desktop

Colleen Coll: https://www.linkedin.com/posts/colleen-coll-b971505_cyberresiliency-commvaultshift-continuousbusiness-activity-7249403692909539330-t_ON/?utm_source=share&utm_medium=member_ios

Stephen Foskett: https://www.linkedin.com/posts/sfoskett_shift-virtual-2024-activity-7246904266865414149-Y0o0/?utm_source=share&utm_medium=member_desktop

Stephen Foskett: https://www.linkedin.com/posts/sfoskett_commvault-buys-clumio-activity-7245147682246135808-I-oL/?utm_source=share&utm_medium=member_desktop

Ethan Banks: https://www.linkedin.com/posts/ethanbanks_shift-virtual-2024-activity-7246586952093720576-KygW

Justin Warren: https://www.linkedin.com/feed/update/urn:li:activity:7247338513820459010

Joep Piscaer: https://www.linkedin.com/feed/update/urn:li:share:7246561014610108417

Joep Piscaer: https://www.linkedin.com/posts/colleen-coll-b971505_commvaultshift-continuousbusiness-cyberresilience-activity-7246671971680182274-vx-D/?utm_source=share&utm_medium=member_desktop

Colleen Coll: https://www.linkedin.com/posts/colleen-coll-b971505_commvaultshift-continuousbusiness-cyberresilience-activity-7246671971680182274-vx-D/?utm_source=share&utm_medium=member_desktop


AI and Cloud Demand a New Approach to Cyber Resilience featuring Commvault – Tech Field Day Podcast SpotlightApple Podcasts | Spotify | Overcast | Amazon Music | YouTube Music | Audio

Learn more about Commvault and their approach to cyber resilience on this episode of the Tech Field Day Podcast.


Gestalt IT and Tech Field Day are part of The Futurum Group.


© Gestalt IT, LLC for Gestalt IT: Commvault Shift 2024 Live Blog

View Details

Erwan James is one of the champions of automation in the data center. He has a long association with the technology behind making the data center a better place for operations teams. I talked to him about Nokia Event Driven Automation (EDA), an exciting new platform that will help this process along.

Erwan talked about three big pillars in this interview that are crucial to the vision for Nokia EDA. The first is abstraction. By abstracting the implementation details of a fabric away from the underlying hardware you can focus more on the intent of what you want to accomplish instead of worrying about things like CLI commands or platform capabilities. Nokia EDA understands the components of the system that it operations so all you need to do is declare what you want to happen and the platform will take care of the rest. Today it supports Nokia SR-Linux but soon it will support even more operating systems.

The second pillar is reliablity. Declaring your intent for the network and actually making it happen are two different things. Half-baked automation implementations usually fail when the network crashes because of bad commands or because the system decides to make changes outside of approved windows. Nokia focused on ensuring that EDA was reliable and predictable so that changes and updates were easy to do but also easy to roll back in the event that something is amiss. With the database of modifications that Nokia tracks you can even rollback device software upgrades!

The final pillar is extensiblity. It’s one thing to build a fabric. It’s another to build a good one. Vendors will all tell you that their approach is the best, even if everyone builds them the same way from standards like EVPN and VXLAN. However, the implementation details do matter once you start trying to build systems that interact with each other. Nokia recognized that having a strong opinion about how to build something is an asset, but so too is being able to relax that opinion to make a data center fabric consumable to the masses.

If you want to learn more about the Nokia EDA platform and how Event Driven Automation can help you, please make sure to check out the Networking Field Day Exclusive with Nokia videos available on the Tech Field Day website


© Gestalt IT, LLC for Gestalt IT: Nokia EDA Under the Hood with Erwan James

View Details

I sat down for a great interview with Michael Bushong of Nokia to talk about why automation seems to have stalled out in enterprises. He had some great insights around this topic. He noted that automation has been around in some form or another for almost two decades and it’s something that everyone seems to do but never really commits to.

The assumption is that most people start small by doing something he calls keystroke removal. By automating easy tasks that feel repetitive people have a sense of accomplishment. However, that’s not where the majority of time accumulates inside of an organization. Instead, the handoffs between people and organizations is where the majority of waiting is occurring. When you also factor in how big the efficiency gap is between on-site IT and cloud-based solutions it leads organizations to shelve their attempts.

The advice? Focus on the big workflows and go deeper than simple scripting. Understand those workflows and ensure that you’re solving the right problems instead of the easy ones. Another great suggestion is using Nokia EDA for your rollout. The platform uses common, familiar tools to overcome adoption hurdles while also allowing you to build out your automation projects on your terms with no need to rearchitect the entire enterprise or hire a lot of new staff.

If you want to learn more about what Nokia EDA can help you accomplish, please make sure to check out the Networking Field Day Exclusive with Nokia videos available on the Tech Field Day website


© Gestalt IT, LLC for Gestalt IT: The Promise of Automation with Michael Bushong of Nokia

View Details

As practical applications of AI are rolled out, they are increasingly being deployed on-premises at scale. We are wrapping up this season of Utilizing Tech with Solidigm focused on AI Data Infrastructure by discussing practical deployment considerations with Ariel Pisetzky, VP of Information Technology and Cyber at Taboola in a discussion with Jeniece Wnorowski and Stephen Foskett. Companies like Taboola are built on data and have been deploying AI-driven applications for years. Generative AI brings new capabilities but is part of a spectrum of solutions that leverage data to produce results for customers. As applications mature, many companies are looking to bring them back on-premises, and this trend will likely accelerate given the cost of AI infrastructure as-a-service offerings. Owned infrastructure can also deliver beyond expected lifespans, representing a potential windfall for businesses that can continue to use deprerciated hardware. This is especially true of large flash drives, which have proven much more reliable than initially predicted. Although it is tempting to buy the biggest, fastest infrastructure to extend the lifespan of equipment, Pisetzky recommends focusing on equipment that is flexible and can be re-purposed in other ways in the future. Server storage is unique in that it is easy to upgrade and replace it in place, even hot-swapping drives, and large lives have a very long lifespan.

Apple Podcasts | Spotify | Overcast | More Audio Links | UtilizingTech.com


AI Recommendation System: Getting behind the Scenes with TaboolaOne of the major trends coming out of the AI industry is algorithmic curation. 80% of what viewers watch on OTT platforms, or readers read on news apps are found through the platforms’ recommendation systems.

A Personal RecommenderAI recommendation system executes on the idea that content-based recommendation is built on. They look at users’ interests and activities over a period of time and produce rows of personalized, hyper-specific recommendations to help them find the next thing they may like.

Publishing houses, streaming platforms, retail e-stores and video sharing websites, all leverage integrated recommendation engines on their platforms. These algorithms process billions of user profiles, identifying their tastes and making content recommendations every day.

This season of Utilizing Tech wraps up with a conversation about AI recommendation system with Taboola and Solidigm. Taboola runs an advertising platform that makes content recommendations to billions of active users daily.

But Taboola is not a name average users on the Internet are familiar with. “We’re not a consumer-facing product. So many people might use our products on a daily basis, but not be aware of it,” says Ariel Pisetzky, VP of information technology and cyber.

“We are a content discovery platform. That means that we reside on many of the publishers that you read on a daily basis and love and receive content from.”

Taboola works in the lower layers of applications, helping publishers and advertisers match their content closely with what customers are looking for.

The platform serves 4 billion webpages, recommending upwards of 40 billion articles and stories to users on those platforms every day. “We provide content recommendations for the next thing that you can read, bringing that content to you without actually knowing who you are, without you logging into our service, and without you providing us any specifics about yourself.”

Taboola does this by corresponding the article text to the readers’ interests and reading habits. The recommendations help discover content they may not have chosen initially.

The LLMs behind this use deep learning and natural language processing to sift through nuanced threads underlying the content, and fit them into taste groups.

“You need to understand as a service provider for publishers, how to recognize the article that we reside on, how to identify the users coming into that specific article, the relevant content, and where that user browsing arc is going to end,” he says.

Taboola, like the rest of the industry, leverages artificial intelligence for this. “We have been taking advantage of LLMs and different generative AI technologies to provide additional tools for editors and advertisers to curate article names, the article itself, and imagery that you might get on your beloved websites.”

On-Prem, the Smarter Option for AIFor Taboola, all of that behind-the-scenes data crunching happens at a private data center.

“Storing these vast amounts of data and processing them and creating value out of them is something that when you own the data, you have a lot of advantages, over putting it somewhere in the cloud,” comments Pisetzky.

At private data centers, there is tremendous opportunity to tune and optimize the infrastructure narrowly for the job at hand.

“Owning all of the compute, data storage and networking has proven to be extremely advantageous for training and inferencing,” he adds.

For instance, Taboola has been able to draw out much more performance from CPUs than they offer out-of-the-box with simple optimization techniques like updating code from the vendor libraries.

Even more can be done. “You have multiple layers of optimization – using CPUs in higher capacity, making sure that all of your GPUs and CPUs are fully utilized.”

These help speed up computation, but also significantly reduce the total cost of ownership (TCO) on the whole.

For a slimmer footprint, Taboola uses NVMe drives. “Today the NVMe interface and SSDs provide so much performance, and when you understand the geometry of the drives and start to think about the specific types of space with different drive geometries that fit in different places, and optimize that for your read size, you suddenly get this boost of performance where you can do so much more with your on-prem hardware and investment in CapEx.”

Pisetzky highlights the value Solidigm SSDs bring to this infrastructure. Solidigm drives, besides being highly reliable, also promise great value for money.

“The drives do not fail beyond the expected MTBF. You’re getting a good bargain for drives that maintain their value beyond their three-year depreciation,” he says.

A big reason for that is high capacity that some of Solidigm’s drives pack. The more terabytes a drive has, the more area it has to spread out the write operations, making it less prone to failure from constant activity.

With AI workloads making CPUs work harder, a good storage solution is all the more important to get high bandwidth and more life out of the drives.

“The beautiful thing with Solidigm drives is the connection to the OS and the tooling that it provides. We can manage the drives remotely through the OS providing us with all the serial numbers and asset management information needed to do this in a responsible way.”

At the edge where infrastructure is limited, inferencing can be a difficult prospect if enterprises do not have servers at the ready in the front-end edge data centers. Having a stack handy allows them to ship out servers on demand to the back-end centers. This way, “if you have a data center automation stack, you can provide storage, where you need it, when you need it.”

Pisetzky advises organizations to focus on having a balanced infrastructure rather than rushing to acquire the most expensive equipment on the shelf. One of the advantages on-premise data centers put on the table is the flexibility to right-size the infrastructure to a t.

“We really love to see how year-over-year, we optimize our use cases for storage, for CPUs and for network, and bring them together to a place where our developers now have so much raw power at their fingertips that they just do not want to go to the cloud for many of the day-to-day operations.”

The rising electricity demands in data centers is another reason to keep the footprint to a minimum, and opt for components that are designed to be optimally energy-efficient.

“When you’re in the cloud, you get the great carbon footprint of the clouds that are in renewables, you get wonderful e-waste management and so on and so forth. When you are on-prem, you need to control your own destiny,” Pisetzky reminds.

While there is no way to cap the dizzying amount of power that accelerators burn, there is certainly an opportunity to push down the envelop with energy-rated storage solutions that underpin the system.

Talking about the specific advantages of Solidigm high-density drives, he says, “When you look at the larger capacities that still provide amazing performance in terms of the level of IOPS, their thermal footprint doesn’t warm up your data center and their energy levels are in use only when you are at full write mode.”

Special thanks to Solidigm for sponsoring this season of Utilizing Tech, and to Taboola for joining the discussion.

Find more about Taboola at their engineering blog. To check’s Solidigm’s SSD portfolio, head over to their website, or catch their presentations from the past AI Field Day event. Keep your eyes peeled for the next episode of Utilizing Tech coming soon on your favorite podcasting app.


Podcast Information:Stephen Foskett is the Organizer of the Tech Field Day Event Series President of the Tech Field Day Business Unit, now part of The Futurum Group. Connect with Stephen on LinkedIn or on X/Twitter and read more on the Gestalt IT website.

Jeniece Wnorowski is the Datacenter Product Marketing Manager at Solidigm. You can connect with Jeniece on LinkedIn and learn more about Solidigm and their AI efforts on their dedicated AI landing page or watch their AI Field Day presentations from the recent event.

Ariel Pisetzky is the VP of Information Technology and Cyber at Taboola. You can connect with Ariel on LinkedIn. Learn more about Taboola by heading to their website.

Learn More about Taboola:* Learn more on their website.


Thank you for listening to Utilizing Tech with Season 7 focusing on AI Data Infrastructure. If you enjoyed this discussion, please subscribe in your favorite podcast application and consider leaving us a rating and a nice review on Apple Podcasts or Spotify. This podcast was brought to you by Solidigm and by Tech Field Day, now part of The Futurum Group. For show notes and more episodes, head to our dedicated Utilizing Tech Website or find us on X/Twitter and Mastodon at Utilizing Tech.


© Gestalt IT, LLC for Gestalt IT: Deploying AI Data Infrastructure in the Datacenter with Ariel Pisetzky of Taboola | Utilizing Tech 07×08

View Details

Discussions about AI are everywhere you turn. The speed at which AI initiatives are emerging makes it difficult to distinguish genuine innovation from mere hype. At the AI Field Day event, the quest was to uncover the reality of use cases and technologies. A panel of very competent delegates posed eager questions and observations around the growing relevance of AI across industries.

The Launch of Private AI and A New Era of Innovation In the wake of its acquisition by Broadcom, VMware is undergoing significant transformations, marking a new era for the company. The VMware Private AI solution is a testament to the company’s ambitions to becoming a key player in the future landscape.

As Frederic Van Haren pointed out, VMware is moving up the food chain from a traditional technology provider to a solution provider. With the launch of the Private AI, VMware by Broadcom steps into a promising phase of innovation and growth.

Before the official rollout, VMware has already engaged 60 customers in an early adoption program, functioning as a private beta testing environment. This initiative is driven by a mix of curiosity and recognition of the need for an AI solution that values privacy and customization.

ScreenshotBenefits of Industry-Specific Private AIAt the core of industry-specific Private AI models lies a deep commitment to enhanced data privacy and security, and the ethical use of data—principles that are paramount in today’s digital landscape. VMware has proactively established an AI Council to address the evolution in AI governance.

“We’ve had governance practices that we put into place. It’s an area where we feel we’re ahead of a lot of our peers in the industry that haven’t even set up that type of governance yet. It’s a work in progress, but there’s a lot happening in this space,” shared Chris Wolf, Global Head of AI and Advanced Services, during his session.

The significant improvement in result accuracy and operational efficiency of industry-specific AI models is another step in the right direction. It makes it possible for organizations to achieve more relevant and precise outcomes by refining the models to process and analyze data pertinent exclusively to particular sectors like healthcare.

This specificity not only boosts the effectiveness of AI applications, but also simplifies the process of tuning the models to better meet the unique needs and challenges of that industry.

The customization of AI models to focus solely on industry-relevant data has the added advantage of requiring fewer computational resources. Smaller, more focused models are less demanding in terms of processing power, which translates to reduced operational costs. For businesses, this means the ability to leverage AI technologies becomes more affordable, enabling a wider adoption across sectors with varying budget constraints.

Self Service Portal for Data Scientists The foundation of VMware’s Private AI offering is VMware Cloud Foundation (VCF), a platform it runs jointly with partners like Intel. VCF is a comprehensive operating model designed to ensure that organizations can build a highly efficient and optimized environment.

Justin Murray’s presentation shed light on the transformative aspect of VMware’s approach: The self-service catalog. He describes this feature as the “Nirvana” of the solution, aimed at empowering data scientists by providing them with the easiest and quickest way to access their necessary tools and platforms.

According to Murray, the essence of this service is to strip away the common complexities that data scientists face, such as navigating through networking issues, or worrying about disk space. The goal is to place the focus squarely on enabling them to get their tools up and running with minimal friction.

ConclusionBy prioritizing privacy, customization, and ease of use, VMware’s Private AI solution is setting a new standard for enterprise AI deployments. As we look forward, the implications of these advancements extend far beyond operational efficiencies, promising a future where AI is integral to solving some of the most pressing challenges faced by industries today.

If you want to know more about VMware Private AI, watch the VMware sessions from the recent AI Field Day event. You can also learn more about VMware’s AI initiatives on their website.


© Gestalt IT, LLC for Gestalt IT: Running Enterprise Industry Specific Private AI with Intel, and VMware by Broadcom

View Details

There are many times we don’t truly realize the power technology wields in our lives. From the cars we drive, to the restaurants we frequent, technology is everywhere. Take AI for example. AI has been game-changing the last few years, transforming operations and ramping up innovation worldwide.

One company in particular, has taken AI to the next level. Nature Fresh Farms, a greenhouse farm in Ontario, is leveraging AI to produce the freshest and sweetest berries in the market.

Nature Fresh Farms scaled and transformed their operations using Intel-powered AI. At the most recent AI Field Day 4 event, they spilled the beans about their AI stack.

Bigger Undertakings for Bigger Yields Nature Fresh Farms started as a 16-acre greenhouse many years ago. It was designed with a data-forward approach using one computer. Over the course of time, the overall operations and the behind-the-scenes IT stack have expanded hand in hand.

Their goal is to grow more crops per meter square, and increase the yield year-over-year, said Keith Bradley, VP of IT and Security.

It is no secret that Intel sets the standard for distinguished computer hardware in the market. So, leveraging Intel’s solutions is a no-brainer, he said.

Starting initially with a 3-node cluster, the team soon realized that more power and efficiency would be required to get to the goal. Upgrading to the Intel Gen3 processors and adding more compute, they knew would be the differentiator, to get AI to generate better results and move on from reactive farming to getting ahead of the weather patterns.

With the stack in order, Nature Fresh Farms taps into data. The company captures data through rows of sensors planted across the greenhouse that monitor soil moisture, temperature density, CO2, vegetation growth and a range of other factors. The data is fed into the primary datacenter at the edge for processing.

The defining factor is a host of AI models, 32 till date, that is used to crunch this data and flesh out insights.

These technical changes have amounted to consistent and impressive increase in yield year-over-year. But more importantly, the ability to leverage CO2 emissions to help stimulate overall plant growth is exactly what Nature Fresh Farms needed to see continued success in their operation.

There are still ways to go as within the greenhouse, manual intervention is still key to performing many of the daily tasks. But the team is optimistic that strategic adjustments like this will get them to where they want to be.

AI for the Future and BeyondAnytime we see powerful tools like AI being used in real-life situations, it sparks excitement and optimism. If organizations can continue to build on what LLMs have started, and utilize more hyperconverged systems within the organization, the possibilities will truly be endless.

It’s little wonder that more businesses are leveraging Intel hardware regularly for these kinds of deployments. The continued growth of both Nature Fresh Farms and Intel is nothing short of spectacular. Watching Intel rise continually to the top of the game and change the way AI is deployed is all the more remarkable.

For more, be sure to check out Nature Fresh Farms’ presentation with Intel from the recent AI Field Day event.


© Gestalt IT, LLC for Gestalt IT: Nature Fresh Farms Leverages Intel and AI to Maximize Yield

View Details

The recent AI Field Day event focused three days on the topic everyone is contemplating – the meteoric growth of generative AI, and how it brings new challenges and opportunities for innovation to infrastructure providers.

Intel hosted an entire day of the event, and brought powerful friends with them, including Google Cloud. The audience was looking forward to hear Google Cloud’s perspective on the AI opportunity, and to say it mildly, presenters, Brandon Royal and Ameer Abbas did not disappoint.

Google Cloud’s Proven AI LeadershipGoogle is one of the foundational players in the AI arena. Their DeepMind team is behind some of the world’s most impactful innovations in the industry. Brandon describes their vision of AI as a complete platform shift for the industry, and likens this moment to the introduction of the Internet and mobile eras.

Generative AI is driving change faster than even these predecessors, powered by sweeping adoption of AI models for business transformation.

Open Software at the Heart of Google Cloud’s StrategyAs the infrastructure provider for 70% of the world’s leading generative AI organizations, Google knows something about the speed of change. It supports the AI revolution through a combination of hardware and software innovation. This starts at the heart of the software platform for which Google Cloud has released Gemma, an open source model developed by DeepMind for delivering GenAI applications.

Gemma is based on Google’s Gemini software which is their in-house stack for AI applications. It has been delivered upon the model of Kubernetes release based on Google’s Borg software. Gemma offers a 2 billion and a 7 billion parameter model with base and instruction tuned versions, and broad support across programming languages. Gemma is on the watchlist for broad adoption, given the known depth of skill of its creators.

ScreenshotGoogle Cloud Taps the Broadest Range of AI SiliconNext is the hardware infrastructure, and Google had a lot to say about platform requirements for both training and inference. Google’s AI services are built around their Kubernetes Engine, and a foundational platform for AI. This platform is quite extensive in terms of flexibility of processor choice as well as scale which they call out as being the highest performing in the industry. They offer a combination of CPUs, NVIDIA GPUs, and Google’s own TPU platforms.

Clear guidance is provided on use of CPUs for low-cost and small to medium model inference, GPUs for medium to large model inference, fine tuning, and medium to large model training, and TPUs for everything from medium model inference to large model training.

The presentation gets deeper into CPU platform capabilities and points to Intel AMX technology for providing fantastic support for applications, including natural language processing, recommendation engines, image recognition, object detection, and media and video analytics.

Intel has invested in the unique acceleration of AI models, and AMX is the latest differentiator for them.

Google is also making a major play on their own TPU technology, as one of the earliest cloud providers to invest in custom-grown silicon. Next generation cloud TPU v5p is on tap for this year with scale to 9K chips with distributed shared memory, driving support for training of the largest AI models and setting Google Cloud apart from competition.

ConclusionBased on the depth of discussion of the heritage and roadmap for TPUs, it’s obvious that Google is making a huge play in differentiated services with the delivery of TPU and Gemma, and hoping to gain market share as generative AI ignites across enterprise verticals. Google remains committed to offering customers a choice, and their CPU and GPU instances demonstrate that they will deliver to customer requirements instead of trying to force-fit everything into a TPU.

Customers are highly likely to respond well to the clear guidance on instance offerings, and hopefully a good customer uptake will follow on Gemma as a core tool in the AI toolbox. While competition from other cloud service providers will be fierce no doubt, the generative AI era will certainly aid in delivering advancements for Google’s service offerings with customers.

Check out Google Cloud’s presentation from the recent AI Field Day event to get in the weeds of their AI infrastructure and solutions.


© Gestalt IT, LLC for Gestalt IT: Google Cloud Re-Architects Infrastructure for AI Era

View Details

In the realm of artificial intelligence (AI), the spotlight often shines on cutting-edge accelerators like GPUs and specialized chips. But, amidst this fervor, a quiet revolution is unfolding within the CPU market. As AI applications grow and evolve at a rapid pace, CPUs are carving out a distinct niche for themselves in the landscape of inferencing.

Industry experts shared insights about this at the AI Field Day event in February, where Intel hosted a full day of presentation. Ro Shah, AI Product Director at Intel, presented a session on the evolving dynamics of AI inference, with a particular focus on the role of CPUs.

ScreenshotA Paradigm Shift in Deployment“When discussing AI, we often traverse the entire spectrum from data processing to model training and deployment,” said Shah. “I’d like to specifically zoom in on the inference phase, where CPUs are increasingly proving their mettle.”

Traditionally, CPUs have been synonymous with data processing tasks, while GPUs have taken the lead in AI model training. Shah’s remark highlights a paradigm shift in the deployment scenarios. Increasingly, CPUs are gaining traction in AI tasks owing to their improved versatility and efficiency gains.

Delving deeper into the nuances of AI inference, and the diverse customer usage models, Shah commented, “We observe a bifurcation in deployment scenarios that range from AI-centric applications with large cycles, to scenarios where a mix of general-purpose and AI cycles converge.”

There are clear-cut reasons for this shift. “Customers are increasingly turning to CPUs for inference due to several key factors,” Shah said. “We consistently hear that CPUs meet critical requirements, enable ease of deployment, and offer compelling total cost of ownership (TCO) benefits for a wide array of workloads, including general-purpose and AI tasks.”

Shaw presented data that shows the improvements in AI workload processing, including advancements in CPU architecture. These datapoints indicate that modern generations of CPUs can handle AI workloads with unprecedented efficiency.

CPU and Accelerator DynamicsWhile there is a growing ecosystem of accelerator alternatives targeting diverse AI applications, Shah highlighted that one must also recognize their limitations in handling extremely large language models.

He outlined the thresholds where CPUs can shine, and tasks where accelerators become indispensable. “For models below 20 billion parameters, CPUs can meet critical latency requirements, offering a compelling choice for many enterprises. Beyond this threshold, accelerators come into play, addressing the burgeoning demand for processing power,” he explained.

It’s evident from the discussion that CPUs, fueled by advancements in architecture and a growing demand for versatile computing solutions, are making a breakthrough in AI. While accelerators will no doubt continue to dominate certain niches, CPUs are poised to rake up a significant market share, bringing to the shelves a compelling alternative for a wide array of AI workloads.

Wrapping UpAccelerators hold undeniable advantages in handling colossal models and specialized tasks. The resurgence of CPUs underscores the importance of versatility and adaptability in AI deployment. As businesses navigate the complexities of AI adoption, a nuanced understanding of the strengths and limitations of both CPUs and accelerators will be the key to unlocking the true potential of AI. The secret lies in striking a balance, leveraging the strengths of what might be the most suitable hardware for the workload on hand.

For more, be sure to watch Intel’s full presentation from the AI Field Day event, and other resources on Gestalt IT.


© Gestalt IT, LLC for Gestalt IT: Unveiling the Role of CPUs in AI Inference and a Growing Trend of Accelerator Alternatives with Intel

View Details

Model training seriously stresses data infrastructure, but preparing that data to be used is a much more difficult challenge. This episode of Utilizing Tech features Subramanian Kartik of VAST Data discussing the broad data pipeline with Jeniece Wnorowski of Solidigm and Stephen Foskett. The first step in building an AI model is collecting, organizing, tagging, and transforming data. Yet this data is spread around the organization in databases, data lakes, and unstructured repositories. The challenge of building a data pipeline is familiar to most businesses, since a similar process is required in analytics, business intelligence, observability, and simulation, but generative AI applications have an insatiable appetite for data. These applications also demand extreme levels of storage performance, and only flash SSDs can meet this demand. A side benefit is the improvements in power consumption and cooling versus hard disk drives, and this is especially true as massive SSDs come to market. Ultimately the success of generative AI will drive greater collection and processing of data on the inferencing side, perhaps at the edge, and this will drive AI data infrastructure further.

Apple Podcasts | Spotify | Overcast | More Audio Links | UtilizingTech.com

Storage, Not an Afterthought in AI – A Conversation with VAST DataOn the floors of the world’s biggest artificial intelligence labs, rows of supercomputers outfitted with powerful workhorse accelerators stand ready to whir into action and start churning data. Except, companies have a problem. It’s the data.

After exhausting all reservoirs of data on the Internet, they are now turning to internal data to train their next version of AI systems. But this is not so much a supply problem as it is a handling problem.

The data is culled and corralled from different sources and silos, prepped and readied before it is exposed to the models. Methodically managing this growing diversity of data assets across the digital gulf has become the Achilles heel of many organizations.

It takes a deep knowledge of data science. The process involves endlessly processing information, analyzing data quality, and decision making to just be able to pick and choose what data to use.

Underneath, a robust and muscly storage system needs to be in place to bear the weight of this proliferating corpus.

In this episode of Utilizing AI Podcast Season 7, Subramanian Kartik, Global Vice President of Systems Engineering at VAST Data, joins co-hosts, Stephen Foskett, and Solidigm’s Datacenter Product Marketing Manager, Jeniece Wnorowski, to talk about this. The discussion shines the spotlight on data, the new celebrity in the AI arena. This episode is brought to you by Solidigm.

The Place of Storage in AI Is Not in the SidelinesGPUs and DPUs often hog all the limelight in conversations about artificial intelligence, and rightly so. “GPUs are important, no question. AI companies need data, and they have a lot of performance requirements due to checkpointing and other things,” says Kartik.

But a fact that is underwritten and often overlooked is that a supporting storage infrastructure equally high-performing and efficient is vital for GPUs to be able to function at max capacity.

“There’s actually a whole bunch of heavy-lifting that happens before gigabytes of tokens are created distilled from petabytes and petabytes of data all on the internet to build these foundation models which we all know and love,” Kartik reminds.

The AI workflow is a highly dynamic one where every piece is different. They have their own characteristics and peculiarities, not to mention widely different moods and needs. The only thing they all have in common is data.

“Data fits as part of the pipeline that spans all the way from raw data to inference, and each of the stages is data-intensive as we go along, not just a little bit under the GPUs.”

With transfer learning picking up, companies are reaching into their own archives for business-specific data to fine-tune private LLMs. Harvesting this data is more trouble than anticipated.

There are decades of data, some archived since before the organizations even existed, that needs to be sifted through and analyzed end-to-end to correctly conclude what datasets will be valuable for AI learning.

“This is a data wrangling problem which is enormous, and customers have tens, if not hundreds of petabytes of data. We’re now trying to figure out how to get a grip on this, and actually make these fit in the old AI pipeline,” Kartik says.

It does not stop there. With data accumulated over many years, there is no one place to find everything. Assets are scattered across databases, data lakes, data warehouses, cloud, on prem, and so on. Tracking down each of these silos, and accessing the data within while getting around the limitations of the dated technology they are sitting on, is a major obstacle in and of itself.

“There is no ontological model which typically exists in large enterprises and people are scrambling to build one. You will hear a lot of talk about things like knowledge graphs, data fabrics and data meshes. These are all efforts to get a grip on where the data is, what is it, and what use is it to them,” Kartik informs.

If a company is able to somehow locate all the datasets that the model can learn from, the next big challenge is to clean and refine that data and make it ready for consumption by the LLM. It will require a heavy-duty storage solution that can not only accommodate that colossal volume of information, but also make it available to the GPUs with speed and consistency.

Ironically, no money is made in training. “Money does not get made in training – it’s a money sink,” he exclaims. “It’s all in the inference. You’ve got to get it to the hands of the end users and that’s what needs to happen.”

A Case for SSDsMany companies have undergone colossal infrastructure overhauls to get AI-ready. “Part of this transformation which we think is going to be the biggest infrastructure transformation in the history over the next five years or so, and we anticipate people will spend about $2.8 trillion on this, is to start aggregating the data and understanding the meaning,” Kartik predicts. “The next thing is to prepare it and decide how they’re going to vector it towards a variety of models which are then going to be able to transform how the business operates.”

In AI and HPC spaces, businesses are rapidly upgrading servers with solid-state technologies to get the system ready for AI workloads.

“The I/O patterns for AI tend to be more random read dominated, and NAND can do a much better job delivering this,” Kartik comments.

Hard disks are falling out of favor, and traditional hard drive-based storage systems are slowly giving way to all-Flash namespaces. “The drives need to handle an unusual mix of workloads – traditional HPC simulation, MPI jobs – high-throughput, large block-type sequential read/write workloads, contrasting with the heavy random I/O intensive workloads. The perfect platform for this combination is a completely solid-state solution. That transformation is well underway,” he says.

The rise of edge and rapid infusion of technology in these environments is hastening this transformation. “From a capacity perspective, floor-space or power, it is going to be significantly lower in these kinds of environments than what we had in disk space environments,” Kartik points out. “I think just that delta alone is going to eliminate hard drives. They’re just not power-efficient enough, not space-efficient enough and not performant enough. They’re going to get squeezed out on the low end with tape, and on the high end with all Flash-based systems.”

Economics is the biggest obstacle. SSDs are significantly pricier, and at scale, well outside budget for many small-size organizations.

Solidigm is changing that reality with its portfolio of high-density, high-endurance solid-state drives. The drives are a balanced combination of performance and capacity, complemented with thin form factors and efficient thermal management.

“Training is a GPU-bound process. That’s not where you get hammered for I/O. Where you do is while doing checkpointing. You absolutely need to have systems that are very high-performance,” emphasizes Kartik.

Naysayers argue that a low volume of data that is typical in small training jobs can lead to idle arrays of expensive drives in datacenters. This argument does not hold for AI workloads because of their performance-intensive nature and unpredictable requirement graph. There’s no way of predicting when you’re going to need more capacity. Having adequate storage at the ready always makes sense for these scenarios.

But this does not necessarily translate to big investments. SSDs like Solidigm’s are built with AI’s most pressing issues – hardware costs and power consumption – top of mind. The drives, being high-density, offer more bang for the buck. They come in slim form factors, meaning more drives can be packed in less space leading to reduced physical footprint and energy consumption.

But what makes them a truly great fit for AI is their ability to maximize GPU utilization. Solidigm SSDs are designed to handle the varying I/O characteristics of the AI phases, and provide high read/write performances throughout. “You no longer have to worry if the right data is at the right place at the right time,” Kartik says.

VAST Data Platform for the AI PipelineWith AI emerging as the new frontier, a new class of cloud service providers (CSPs) have emerged, that, unlike the biggest CPSs we know, are built from ground up to support AI use cases. These companies have the most coveted hardware purpose-built for AI’s scale.

CoreWeave, a specialized CSP that is involved in massive-scale AI training is one of the names in that space. CoreWeave’s secret config is a combination of Solidigm’s QLC SSDs and the VAST Data solution. With these, it has designed a platform that offers the perfect balance of scale, speed, performance, and efficiency in the AI data pipeline.

VAST Data plays a crucial role in untangling the AI data knot. The VAST Data platform exposes data through modalities other than the usual file and object protocols which makes the lengthy process of AI data management less costly and cumbersome.

“We also expose ourselves through tables,” explains Kartik. “The native tabular structures within us is crucial for data crunching – taking the raw data which is currently sitting in large data lakes on Hadoop, Iceberg, MinIO or object stores. We want to be able to corral that and give it the transformation platform to convert into what can be actually put into a model.”

The platform caters to all the stages of the AI pipeline – data capture, data preparation, training and inferencing. “We do this with an exceptional degree of security, governance and control.”

Consolidating the entire pipeline into a single platform eliminates the need to copy and move data around, significantly reducing data footprint. This unification, Kartik says, is what makes VAST Data a highly suitable solution for any AI work.

You can now listen to Utilizing AI in your favorite podcast application. Be sure to give this episode a listen, and keep your eyes peeled for upcoming episodes. To read up more on this, check out VAST Data’s white papers on their website. To know more about Solidigm’s high-performance, value SSD solutions for AI, head over to their website.

Podcast Information:Stephen Foskett is the Organizer of the Tech Field Day Event Series President of the Tech Field Day Business Unit, now part of The Futurum Group. Connect with Stephen on LinkedIn or on X/Twitter and read more on the Gestalt IT website.

Jeniece Wnorowski is the Datacenter Product Marketing Manager at Solidigm. You can connect with Jeniece on LinkedIn and learn more about Solidigm and their AI efforts on their dedicated AI landing page or watch their AI Field Day presentations from the recent event.

Subramanian Kartik, Ph. D, is the Global Systems Engineering Lead at VAST Data. You can connect with Subramanian on LinkedIn and learn more about VAST Data on their website or watch the videos from their recent Tech Field Day Showcase.

VAST Data Tech Field Day Showcase:* Running Full Stack AI Operations at Scale with VAST Data and Run:ai * VAST Data DASE Architecture Optimized for Supermicro Hyperscale * Running VAST Data End to End on NVIDIA BlueField DPUs * Operationalizing AI at Scale with VAST Data


Thank you for listening to Utilizing Tech with Season 7 focusing on AI Data Infrastructure. If you enjoyed this discussion, please subscribe in your favorite podcast application and consider leaving us a rating and a nice review on Apple Podcasts or Spotify. This podcast was brought to you by Solidigm and by Tech Field Day, now part of The Futurum Group. For show notes and more episodes, head to our dedicated Utilizing Tech Website or find us on X/Twitter and Mastodon at Utilizing Tech.


© Gestalt IT, LLC for Gestalt IT: Building an AI Training Data Pipeline with VAST Data | Utilizing Tech 07×02

View Details

Join Tech Field Day and Qlik at Qlik Connect for the Day Two Keynote with invaluable insights, best practices, and insider tips to elevate your data and analytics expertise! Discover Qlik’s vision, product roadmap, and future outlook. Hear from Qlik executives, partners, customers, and industry insiders as they share best practices and innovative applications for data, analytics, and AI, and learn from the former director of the James Webb Space Telescope.

Register now to watch the Keynote live or keep it here to learn what the Tech Field Day Delegates have to say here on our live blog and social queue. The posts below will be ordered from newest to oldest.


Day Two

I was very impressed by the insight from the AI Council! It says a lot that @Qlik is actively seeking guidance from these thought leaders! #QlikConnect https://t.co/n4rhLpkSz5

— Stephen Foskett (@SFoskett) June 5, 2024

https://www.linkedin.com/posts/jaycuthrell_jwst-qlikconnect-space-activity-7204119168789426176-RvTQ/?utm_source=share&utm_medium=member_desktop

Jay Cuthrell

Screenshot

The sun shield ?? on Webb has 1.2 million SPF rating!

— Ben Young (@benyoungnz) June 5, 2024

Rings around Uranus – until Webb, we had never seen these pic.twitter.com/gTOuEJ2T2g

— Ben Young (@benyoungnz) June 5, 2024

Webb project was burning around $1.5 million dollars a day

— Ben Young (@benyoungnz) June 5, 2024

Audience engaged.

Former director of the James Webb program Greg Robinson giving us some great insights into how it works, its purpose and a lot of interesting snippets along the way#QlikConnect pic.twitter.com/A4nnfh3TRD

— Ben Young (@benyoungnz) June 5, 2024

https://www.linkedin.com/posts/sfoskett_ai-qlikconnect-activity-7204112563112923137–vNQ/?utm_source=share&utm_medium=member_desktop

Stephen Foskett

ScreenshotGina Rosenthal on Mastodon

@gminks@mas.to

Screenshot

Now AWS, Accenture, and Qlik are taking the stage for the next panel at #QlikConnect.

— Denny Cherry – mrdenny@techhub.social (@mrdenny) June 5, 2024

The @qlik AI Council on successful AI projects

Treat your data as the first priority.

Most breakthroughs in AI to date have been down to finding / developing the right dataset.#QlikConnect

— Ben Young (@benyoungnz) June 5, 2024

Gina Rosenthal on Mastodon

@gminks@mas.to

Screenshothttps://www.linkedin.com/posts/fredericvharen_qlikconnect-tfdx-activity-7204104307518914560-3W6e/?utm_source=share&utm_medium=member_desktop

Frederic Van Haren

Screenshothttps://www.linkedin.com/posts/sfoskett_qlikconnect-ai-activity-7204101665946255362-quzY/?utm_source=share&utm_medium=member_desktop

Stephen Foskett

Screenshot

This morning's keynote at #QlikConnect begins with a deep and thoughtful conversation about #AI with the new @Qlik AI Council.
/1 pic.twitter.com/bCTsAQIX3W

— Stephen Foskett (@SFoskett) June 5, 2024

It's very easy for corporations to blithely ignore the risks of technology, but Qlik is inviting this conversation to the stage, with Nina Schick and Dr. Rumman Chowdhury pointing out the risks of deepfakes and generated images to individuals and governments. #QlikConnect /2

— Stephen Foskett (@SFoskett) June 5, 2024

It's important to have these voices front and center, because these risks are real and are impacting everyone.#QlikConnect /3

— Stephen Foskett (@SFoskett) June 5, 2024

Michael Bronstein points out that the AI field is advancing incredibly quickly, making it difficult to predict what comes next. But AI has helped in science, including protein folding and biotech data analysis. #QlikConnect /4

— Stephen Foskett (@SFoskett) June 5, 2024

Although most of the conversation will likely focus on risks, and Dr. Bronstein certainly will discuss these, it's important to consider the positives as well.#QlikConnect /5

— Stephen Foskett (@SFoskett) June 5, 2024

Returning to risks of #AI, Kelly Forbes points out that AI makes cyberattacks worse, and industry and government must approach these differently. Dr. Chowdhury points out that traditional statistical models are not prepared to deal with bulk AI-generated data.#QlikConnect /6

— Stephen Foskett (@SFoskett) June 5, 2024

It's a matter of scale and efficiency. We've always had social engineering but AI makes it practical at scale for small groups. James Fisher points out that it's time for companies to develop a strategy to deal with these changes in every area of the business.#QlikConnect /fin

— Stephen Foskett (@SFoskett) June 5, 2024

Gina Rosenthal on Mastodon

@gminks@mas.to

ScreenshotGina Rosenthal on Mastodon

@gminks@mas.to

Screenshot

The @Qlik AI Council on the Relationship between Innovation and Responsibility

"Brakes help drive you faster"#QlikConnect

— Ben Young (@benyoungnz) June 5, 2024

Gina Rosenthal on Mastodon

@gminks@mas.to

Screenshothttps://www.linkedin.com/posts/fredericvharen_qlikconnect-tfdx-activity-7204096305105661953-TTSk/?utm_source=share&utm_medium=member_desktop

Frederic Van Haren

Screenshot

This week, @SFoskett and the @TechFieldDay delegates are at #QlikConnect. During the event, Qlik announced Qlik Answers, their #AI-powered expert response system. Stay tuned for more coverage from the event. #Shorts #TFDx #QlikAnswers pic.twitter.com/2D570d8ekd

— Gestalt IT (@GestaltIT) June 5, 2024

Talking about AI responsibility at #QlikConnect with the Qlik AI pannel right now.

— Denny Cherry – mrdenny@techhub.social (@mrdenny) June 5, 2024

Ready for the Day 2 Keynote at #QlikConnect! pic.twitter.com/zPc5loAbRQ

— Matt Garvin (@MrMattGarvin) June 5, 2024

Gina Rosenthal on Mastodon

@gminks@mas.to

Screenshot


After keeping up with our live blog, check out the Tech Field Day Podcast episode featuring Qlik as well as other episodes featuring our delegates. Thank you for listening or watching!

Apple Podcasts | Spotify | Overcast | Amazon Music | YouTube Music | Audio


© Gestalt IT, LLC for Gestalt IT: Qlik Connect Keynote Live Blog – Day Two

View Details

Join Tech Field Day and Qlik at Qlik Connect for an immersive two-day event packed with invaluable insights, best practices, and insider tips to elevate your data and analytics expertise! Discover Qlik’s vision, product roadmap, and future outlook. Hear from Qlik executives, partners, customers, and industry insiders as they share best practices and innovative applications for data, analytics, and AI, and learn from the former director of the James Webb Space Telescope.

Register now to watch the Keynote live or keep it here to learn what the Tech Field Day Delegates have to say here on our live blog and social queue. The posts below will be ordered from newest to oldest.


Day One

AI mention counter at the #QlikConnect general session

Good persistence @jdanton pic.twitter.com/2J9VYgx5xN

— Ben Young (@benyoungnz) June 4, 2024

Jay Cuthrell on Mastodon

@jay@cuthrell.com

Screenshot

Qlik Answers will allow companies to create KBs from their existing unstructured data and use those to get answers from their AI platform.

#QlikConnect

— Denny Cherry – mrdenny@techhub.social (@mrdenny) June 4, 2024

https://www.linkedin.com/posts/jaycuthrell_passionatepractitioners-qlikconnect-activity-7203758627424272385-45Vm/?utm_source=share&utm_medium=member_desktop

Jay Cuthrell

Screenshot

At scale, RAG is very difficult if you have ever tried to build a stack.

Qlik Answers has built a complete experience that has indexing, vector database, model serving, user interface (embeddable) and monitoring.

All the guess work around chunking strategies, system prompts… pic.twitter.com/6ylElXYTx3

— Ben Young (@benyoungnz) June 4, 2024

Gina Rosenthal on Mastodon

@gminks@mas.to

Screenshothttps://www.linkedin.com/posts/fredericvharen_qlikconnect-tfdx-activity-7203758068214419457-qcXh/?utm_source=share&utm_medium=member_desktop

Frederic Van Haren

Screenshot

Creating a new app on trusted data from an organisations own data marketplace.

Supply chain data being consumed to create some new reports as an analytics analyst. #QlikConnect pic.twitter.com/y8OFVSBXyF

— Ben Young (@benyoungnz) June 4, 2024

I am very impressed by how @Qlik is bringing Talend technology to their customers. Read more about Qlik Talend Cloud in this LinkedIn Pulse article! #QlikConnect #QlikAmbassador https://t.co/VqJAGC4EN2

— Stephen Foskett (@SFoskett) June 4, 2024

https://www.linkedin.com/pulse/qlik-cloud-integration-talend-stephen-foskett-n0pze

Stephen Foskett – I wrote a quick LinkedIn Pulse article about Qlik Talend Cloud!

Screenshot

This structure reminds me a lot of what Jeff Bezos mandated back at Amazon years ago where all interactions between divisions must be over APIs and service contracts.

Each division provides a set of services others can consume if they need. #QlikConnect

— Ben Young (@benyoungnz) June 4, 2024

This is a new modern organisational structure where we break down the traditional IT / Business department silos by having data products and owners per division

Divisions creat “Data Products”#QlikConnect pic.twitter.com/YH9NfjiSdX

— Ben Young (@benyoungnz) June 4, 2024

Qlik showing their Data Products feature, built on the concept of data mesh. Some nice data governance and quality features are built in here. #qlikconnect #tfdx pic.twitter.com/ZE7LyLHckf

— Joey D'Antoni (@jdanton) June 4, 2024

https://www.linkedin.com/posts/fredericvharen_qlikconnect-tfdx-activity-7203750372841009154-S9Dr/?utm_source=share&utm_medium=member_desktop

Frederic Van Haren

Screenshot

When all of this comes together, this is what the Qlik Trust Score for AI looks like

See changes over time, monitor and react as needed ?#QlikConnect pic.twitter.com/F91A5PrVU1

— Ben Young (@benyoungnz) June 4, 2024

Gina Rosenthal on Mastodon

@gminks@mas.to

Screenshot

Qlik has built a Trust Score for AI and it is calculated on the following dimensions:

Diverse (siloed data creates bias results)

Secure (data will have PII, financial etc, make sure machines are not leaking data. Mapping data and masking, access management and role mapping)…

— Ben Young (@benyoungnz) June 4, 2024

Live demo time at #QlikConnect

Starting with data ingest in Qlik Cloud pic.twitter.com/POZJ8GKgU8

— Ben Young (@benyoungnz) June 4, 2024

Gina Rosenthal on Mastodon

@gminks@mas.to

ScreenshotScreenshot

Qliks model builds a trust score for their AI models. #qlikconnect #tfdx pic.twitter.com/rJ0R2nFR9q

— Joey D'Antoni (@jdanton) June 4, 2024

Demo of Qlik Cloud ETL moving data between a wide variety of sources #tfdx #qlikconnect pic.twitter.com/DDctoqkvTu

— Joey D'Antoni (@jdanton) June 4, 2024

The generation of SQL from plain English that they are showing at #QlikConnect is a pretty simple solution so that business people, not IT people, can build data transformations.

— Denny Cherry – mrdenny@techhub.social (@mrdenny) June 4, 2024

Qlik's AI messaging and products are more data focused than a lot of other messaging I've seen, and it's refreshing. #QlikConnect #TFDX

— Joey D'Antoni (@jdanton) June 4, 2024

Time for. Demo of Qlik Cloud at #QlikConnect. Their UI is a pretty slick experience for moving data around. pic.twitter.com/VtVcsi8oDn

— Denny Cherry – mrdenny@techhub.social (@mrdenny) June 4, 2024

Qlik and AWS have announced a strategic collaboration agreement.

Co-Innovation on new AI solutions
Marketing/sales collaboration

Over 7k customers using Qlik on AWS#QlikConnect

— Ben Young (@benyoungnz) June 4, 2024

https://www.linkedin.com/posts/sfoskett_qlikconnect-ai-sap-activity-7203742138772172800-pdAA/?utm_source=share&utm_medium=member_desktop

Stephen Foskett

Screenshot

Nina Schick is on stage at #QlikConnect with bold statements! She said geopolitics demand sovereign AI, and that trust, transparency, authentication, and ethics when it comes to data-driven AI. Tomorrow's keynote will feature the @Qlik AI Council deep dive into these topics. pic.twitter.com/G5BJmhJ5bV

— Stephen Foskett (@SFoskett) June 4, 2024

Sovereign AI will be the core to success in the future of business. #QlikConnect

— Denny Cherry – mrdenny@techhub.social (@mrdenny) June 4, 2024

You can be good at AI, or you can be bad at business.
–Nine Schick #QlikConnect

— Denny Cherry – mrdenny@techhub.social (@mrdenny) June 4, 2024

Gina Rosenthal on Mastodon

@gminks@mas.to

ScreenshotScreenshot

Such an important function for organisations in ensuring we are responsible with AI

Qlik announcing their AI Council#QlikConnect pic.twitter.com/P5b1kKutQB

— Ben Young (@benyoungnz) June 4, 2024

Build or buy AI models. IDC it's construct, or consume.

Fastest way to get access to GenAI is to lean on application vendors
Fine tuning in some circumstances (higher cost, complexity)
And in even more bespoke situations, train (high cost)#QlikConnect

— Ben Young (@benyoungnz) June 4, 2024

IDC Research showing what data sources will be leveraged by organisations in GenAI applications

Customer data – 82%
Internal data – 72%
Public / open source data – 67%
Licensed data – 58%#QlikConnect

— Ben Young (@benyoungnz) June 4, 2024

"Build vs. Buy" your AI models being discussed on the KeyNote stage right now at #QlikConnect

— Denny Cherry – mrdenny@techhub.social (@mrdenny) June 4, 2024

https://www.linkedin.com/posts/fredericvharen_qlikconnect-activity-7203736751813578753-Mnzp/?utm_source=share&utm_medium=member_desktop

Frederic Van Haren

Screenshothttps://www.linkedin.com/pulse/qlik-connect-keynote-joseph-d-antoni-msbze/?trackingId=%2BmD85v4aoMU9Jr6WAqZM3g%3D%3D

Joey D’Antoni

Screenshot@gminks@mas.to

Gina Rosenthal on Mastodon

New product announcements here at #QlikConnect

Qlik Talend Cloud – modern code/no code platform to create a trusted foundation for AI analytics and operations

Qlik Answers -GenAI over your unstructured data

— Ben Young (@benyoungnz) June 4, 2024

Get the data foundation right, then add AI #QlikConnect

Data is powerful. ?

Vale – Saved $600M in ONE year using Qlik
Urban Outfitters – using it in 650+stores globally
Penske – using it to perform 90k preventative maintenance operations every year

— Ben Young (@benyoungnz) June 4, 2024

https://www.linkedin.com/posts/sfoskett_qlikconnect-activity-7203737615659843585-V2pd/?utm_source=share&utm_medium=member_desktop

Stephen Foskett

Screenshot

Qlik Talend Cloud is here for use to make BI proceeded quicker and easier. #QlikConnect pic.twitter.com/oWFRmArIa0

— Denny Cherry – mrdenny@techhub.social (@mrdenny) June 4, 2024

https://www.linkedin.com/posts/fredericvharen_qlikconnect-tfdx-activity-7203735428200345600-FHZn/?utm_source=share&utm_medium=member_desktop

Frederic Van Haren

Screenshot

Qlik Talend Cloud is here for use to make BI proceeded quicker and easier. #QlikConnect pic.twitter.com/oWFRmArIa0

— Denny Cherry – mrdenny@techhub.social (@mrdenny) June 4, 2024

Some big wins for Qlik customers over the last few years. Those are some impressive wins. #QlikConnect pic.twitter.com/WO7lYigloi

— Denny Cherry – mrdenny@techhub.social (@mrdenny) June 4, 2024

Great opening message here from CEO @MikeCapone at #QlikConnect

Nice to see the different angles of AI talked about in this pivotal moment in our history. #tfdx pic.twitter.com/0ORiJdIPib

— Ben Young (@benyoungnz) June 4, 2024

https://www.linkedin.com/posts/sfoskett_qlikconnect-ai-activity-7203735106740465665-90ce/?utm_source=share&utm_medium=member_desktop

Stephen Foskett

Screenshot

Mike is taking the stage at #QlikConnect for the keynote. pic.twitter.com/Dt4aw1bvP0

— Denny Cherry – mrdenny@techhub.social (@mrdenny) June 4, 2024

https://www.linkedin.com/posts/fredericvharen_qlikconnect-data-tfdx-activity-7203733459150761985-4v9z/?utm_source=share&utm_medium=member_desktop

Frederic Van Haren

Screenshothttps://www.linkedin.com/posts/gminks_qlikconnect-activity-7203731511064338432-8_7s/?utm_source=share&utm_medium=member_desktop

Gina Rosenthal

Screenshothttps://www.linkedin.com/posts/fredericvharen_qlikconnect-tfdx-activity-7203732259235868672-_TnI/?utm_source=share&utm_medium=member_desktop

Frederic Van Haren

Screenshot

The one and only @KatieLinendoll kicking things off here at #QlikConnect pic.twitter.com/XNyISXTRHt

— Ben Young (@benyoungnz) June 4, 2024

https://www.linkedin.com/posts/sfoskett_qlikconnect-data-activity-7203448141004558337-PG0e/?utm_source=share&utm_medium=member_desktop

Stephen Foskett

Screenshot

Getting ready for #QlikConnect in Orlando. Can't wait to see what they have to show us today. pic.twitter.com/EGsUld1lx7

— Denny Cherry – mrdenny@techhub.social (@mrdenny) June 4, 2024

https://www.linkedin.com/posts/fredericvharen_qlikconnect-ai-data-activity-7203727998624149506-oXCs/?utm_source=share&utm_medium=member_desktop

Frederic Van Haren

Screenshothttps://www.linkedin.com/posts/jaycuthrell_qlikconnect-activity-7203730822309298177-6eoW/?utm_source=share&utm_medium=member_desktop

Jay Cuthrell

Screenshothttps://www.linkedin.com/posts/gminks_qlikconnect-activity-7203368359474696193-YCPV/?utm_source=share&utm_medium=member_desktop

Gina Rosenthal

Screenshot

Brings a tear to the eye #QlikConnect

I need know the history on this though, was there a controversy last time?! ? pic.twitter.com/oCV5ci2U58

— Ben Young (@benyoungnz) June 4, 2024

It feels like a hot minute since I was at a keynote – looking good #QlikConnect

Kudos to the song selection while we wait for things to kick off ? pic.twitter.com/hv8VufQ5Da

— Ben Young (@benyoungnz) June 4, 2024


After keeping up with our live blog, check out the Tech Field Day Podcast episode featuring Qlik as well as other episodes featuring our delegates. Thank you for listening or watching!

Apple Podcasts | Spotify | Overcast | Amazon Music | YouTube Music | Audio


© Gestalt IT, LLC for Gestalt IT: Qlik Connect Keynote Live Blog – Day One

View Details

As artificial intelligence is rapidly advanced and deployed, the infrastructure supporting training and reference data has emerged as a critical foundation for success. Advanced storage solutions are required to support the intensive demands of AI applications for speed, efficiency, and scalability. This gives rise to the concept of “AI Data Infrastructure,” storage and data platforms required to support AI applications. Throughout 2024 we are exploring the relationship between AI and storage, from a series of articles following AI Field Day through a season of the weekly Utilizing Tech podcast through the summer, to a Tech Field Day event focused on AI Data Infrastructure in October. We are learning about AI data infrastructure from our peers, key companies like Solidigm and their partners, and end users deploying modern AI solutions.

The Crucial Role of Storage in AIAI applications are notorious for their extreme I/O demands, requiring robust and responsive storage solutions to function optimally. Jim Czuprynski in his article highlights the complexity of AI workloads and their storage needs, particularly emphasizing the economic impact of underutilized resources: “An idle GPU is one of the highest infrastructure debts in the enterprise AI computing landscape.” This truism came up throughout season 6 of Utilizing Tech as well, where the conversation frequently returned the need for storage that not only keeps GPUs continuously busy but also handles different data access patterns efficiently.

The performance of AI systems is directly linked to the efficacy of the underlying storage technology. Efficient storage systems ensure that AI models operate without interruption, a critical factor in environments where downtime can mean significant financial losses. Jeffrey Powers notes that “storage needs to run efficiently, and at speed, so that GenAI models can run optimally without stopping even for a second.” The emphasis on speed and efficiency reflects the direct impact of storage performance on the overall effectiveness and reliability of AI applications.

Advancements in storage technology, particularly through companies like Solidigm, have been pivotal in meeting these AI demands. As discussed in Ben Young’s article, NAND flash storage devices (SSDs) are preferred due to their large capacities and efficiency advantages over traditional hard disk drives: “Big storage capacity ensures that there is enough space to house the burgeoning datasets that are utilized in the AI workflow.” Expansive and scalable storage solutions are required to accommodate the growing data requirements of AI systems.

Integration of AI Data InfrastructureThe management of AI workloads requires a structured approach to data handling, which is facilitated by advanced storage architectures. Andy Banta details how Supermicro and Solidigm address these requirements through a tiered system architecture that optimizes data flow from ingestion to processing and long-term storage. This structured approach ensures that each phase of the AI workflow is supported by appropriate storage solutions, enhancing the efficiency and speed of AI operations.

Sustainability is an important aspect of modern IT infrastructure, and this is especially true of power-hungry AI deployments. The partnership between Supermicro and Solidigm focuses on creating environmentally friendly datacenter solutions. These solutions not only meet the demanding requirements of AI workloads but also significantly reduce energy consumption and physical space requirements, aligning with broader environmental goals.

As Colleen Coll discussed, the relationship between AI development and storage technology is complex. As demonstrated throughout Solidigm’s presentation at AI Field Day, each phase of the AI pipeline (from data ingestion to inference) has specific storage demands. Storage isn’t just a repository but a dynamic component that must be tailored to support the intensive and varied data handling requirements of AI systems.

Diving Deeper Into AI Data InfrastructureLooking ahead, we will continue to explore the role of AI data infrastructure through Season 7 of Utilizing Tech as well as at our upcoming Tech Field Day events. As AI applications grow more complex and data-intensive, the need for innovative storage solutions that can handle massive datasets, ensure high-speed data access, and provide reliable and sustainable operations will only intensify. The insights from the Utilizing Tech podcast and the contributions of Solidigm and their partners will be crucial in navigating these challenges.

The integration of advanced storage solutions into AI data infrastructure is fundamental to the success and scalability of AI technologies. Efficient, scalable, and sustainable storage solutions are not merely supportive elements but are central to the operational excellence and innovative potential of AI systems. As AI continues to transform industries, the evolution of AI data infrastructure will play a pivotal role in enabling these technologies to reach their full potential, driving forward the next generation of technological advancements.

For more on Solidigm and their products, head over to their website. You can watch their AI Field Day presentations on the Tech Field Day website. Solidigm is also featured on our current season of Utilizing Tech and will cohost our next season beginning June 6, 2024.


© Gestalt IT, LLC for Gestalt IT: The Critical Role of Advanced Storage in AI Data Infrastructure

View Details

In a rapidly evolving tech landscape, the intersection of AI and storage is becoming increasingly pivotal. The storage infrastructure plays an outsize role in optimizing AI workflows. From efficient checkpointing mechanisms to scalability, to high throughput, a good storage solution provides all the imperatives to tap into artificial intelligence.

To delve into this intricate overlap of the two domains, Roger Corell, Director of Solutions Marketing and Corporate Communications at Solidigm, and Subramanian Kartik, Global VP of Systems Engineering, for VAST Data, sit down for a chat. Their conversation explores the multifaceted phases of the AI pipeline, and the corresponding storage demands.

The Five Phases in an AI Pipeline“Storage plays a pivotal role in the efficiency and performance of AI workflows, encompassing the diverse phases from data ingestion to inference,” comments Kartik.

There are five distinct phases in the AI workflow –

  • Data Ingestion: It is a critical process where raw data is pulled from multiple sources for processing and analysis. This phase involves heavy inbound writes. Data is typically sourced from cloud repositories or internal databases.

“Efficient handling of large datasets is crucial to ensuring smooth operations,” says Kartik.

  • Data Preparation: In the second step, the data is cleansed and normalized before being pumped into the model. “Data tends to be dirty and needs to go through normalization of some kind,” Kartik explains.

It is predominantly CPU-bound, and despite substantial reading, the volume of data written back is significantly less at this stage.

  • Model Training: This where the learning begin. Model training necessitates random IO patterns to prevent model memorization. Optimal batch sizes and GPU utilization are key factors influencing model convergence.
  • Model Checkpointing: The neural network’s weights and other parameters need to be preserved periodically as training continues so that it can be rolled back to a previous state if anything goes sideways. This phase involves large block sequential writes. Efficient IO operations are critical in minimizing downtime and data loss. The caveat though, is, checkpointing doesn’t always work. “Bad things happen,” says Kartik.
  • Inference: If everything goes right, at this point, the model should be able to make inferences from data. Primarily IO-bound, inference demands high-throughput reads, especially with larger datasets. Efficient storage performance is essential, particularly for tasks like image generation.

The significance of a robust storage solution in optimizing these phases is undebatable. Notably, Solidigm’s architecture provides certain unique advantageous in handling AI workloads’ asymmetric demands. Designed to support data and read-intensive workloads, its superior read capabilities, high capacity, endurance, and overall efficiency, are key to enhancing processes like checkpointing and inferencing.

ScreenshotSSDs for AI Data StorageUnderstanding and addressing the storage challenges inherent in AI pipelines is critical for organizations in the frontline of innovation. At the AI Field Day event in February, Ace Stryker and Alan Bumgarner built on Kartik’s point in their presentation that underlines the need for a robust storage to handle AI’s data explosion.

Their presentation highlights SSDs as game-changers for AI workflows – from data ingestion to real-time inferencing – echoing Kartik’s insights on scalable architectures and efficient checkpointing.

ScreenshotWrapping UpIn the realm of AI-driven transformation, Kartik’s breakdown paints a clear picture of how computational power and storage go hand in hand. Each AI phase has a mix of hurdles and opportunities, which underscore just how crucial storage is in the making or breaking of the workflows.

Businesses need top-notch storage that can handle this flood of info. Whether it’s gathering data, or making sense of it all, AI’s success hinges on storage that’s tough and flexible. And as deployments get more intricate, companies need to rethink how they approach storage altogether.

Give the conversation a listen at Solidigm’s website. For a technical deep-dive, be sure to watch Solidigm’s presentation from the recent AI Field Day event.


© Gestalt IT, LLC for Gestalt IT: The Intricate Relationship of AI and Storage

View Details

AI’s complex workloads require extreme computing, the kind that only the fastest accelerators are known to provide. It is little surprise that GPUs (Graphics Processing Units) have emerged as the holy grail of compute as AI gains ground. But exorbitant pricing and scarce access, not to mention heavy power consumption and cooling requirements, raise barriers for enterprise adoption.

What if IT shops truly didn’t need these silicon beasts for the bulk of their AI works?

An Affordable Alternative for AI Workloads below 20B ParametersIntel has closely followed the GPU trend over the past few years. To gauge the depth of need of GPUs for AI workloads, the teams have run a series of trials, and they’ve arrived at an interesting conclusion. Based on their findings, CPUs can accommodate almost all AI workloads, with the exception of the insanely intense ones.

For example, the most intense large language model (LLM)-based AI workloads, like Meta’s LLAMA2, typically fluctuate within the range of 7 and 30 billion.

Figure 1. The Reality of AI: GPUs Are Needed Only For Insanely Intense Workloads Intel found that a majority of AI workloads remain below 20 billion parameters. The Xeon chipsets meet almost all latency requirements for the general-purpose workloads in that category. There is rarely the need to leverage the massive acceleration of the GPU technology for these AI workloads, Intel says.

Intel shared benchmarks from real-world scenarios to back this up. In one example of an inference-heavy AI implementation, a customer used Intel Xeon CPUs to perform extremely rapid image processing. The Xeon CPU family was able to scale from a not-insignificant scanning speed of 400 frames per second (fps) using 24 AVX2 CPUs, to over 19,000 fps using 64 cores of their AMX-powered Emerald Rapids processors.

The test results revealed that the newest AMX chipsets engineered with power efficiency kept the 64-core configuration’s energy consumption even with the 24 AVX2.

Figure 2. Same Wattage, Dramatically Increased Workloads: 24 AVX2 Cores vs. 64 AMX Cores Another real-world use case Intel shared is of a customer adding speech translation and real-time transcription services to an existing video conferencing offering. Intel engineers put together a solution with just a few additional servers using Intel CPUs.

Figure 3. Handling LLM Demands: Response Time Latencies Kept Below 100 ms (1/10th of a second)This configuration was particularly interesting because it had to deal with two different AI workloads. Imagine translating the phrase “Pleased to meet you, sir!” to Spanish. The latency for the first word returned in the sequence – “Mucho” – tends to be compute-bound because the AI model needs to find exactly the right word. However, retrieving each of the next words that would be returned in the phrase may also depend on contexts, like the formality of the greeting (e.g. ?Mucho gusto! versus ?Mucho gusto, a sus ordenes!) and tends to be a memory-bound operation.

Intel’s hardware solutions work well for complex AI workloads like the above. Intel has heavily invested in native software support for OSS solutions for data analytics, like Pandas, NumPy, and Apache Spark. Likewise, their commitment to support popular machine learning and deep learning toolsets like PyTorch, TensorFlow, and AutoML (Figure 4) go a long way to extend support to the userbase.

Figure 4. It’s Not Just the Hardware: Intel’s Array of Software Support For Analytics and MLTo GPU, Or Not To GPU?Intel positions its current array of Xeon CPU chipsets for organizations that are grappling with the pressures of providing adequate compute power for AI workloads. Only the most intense AI workloads – specifically, those north of 20 billion parameters – require GPU-level computing power. For everything below that, Intel’s offerings appear to fill the bill. Additionally, their deep support for software compatible with common analytics, machine learning, and deep learning requirements, make their hardware a compelling choice.

Be sure to check out Intel’s presentations on CPUs for AI workloads from the recent AI Field Day event to get a technical deep-dive.


© Gestalt IT, LLC for Gestalt IT: Are GPUs Essential for Every AI Problem? Intel Says, No.

View Details

Artificial intelligence workloads are taxing traditional datacenter infrastructure in unexpected ways. Their need for faster compute, denser memory, and more direct access to GPUs is breaking the limits of what most platforms are capable of offering.

Supermicro, in collaboration with Solidigm, is taking on this challenge. Their mission is to deliver single-sourced, highly efficient, rack-sized storage solutions for AI workloads. March of this year, they made a joint presentation at the AI Field Day event in California, where Wendell Wenjen and Paul McLeod, elaborated on how they are maximizing storage density in datacenter chassis.

What AI Workloads Look LikeThere is a variety of workloads that falls into a single AI workflow. It starts with data ingestion in which raw data is fetched from a variety of different sources. This includes generated data, media content, telemetry from an array of sensors, and so on. This data can be characterized as unstructured, and growing over time. A scale-out object storage solution is ideal for gathering this data.

ScreenshotIn the next step, data is cleaned and transformed to make it usable. This is often a distributed and iterative process. The storage needed for this is high-speed, indexable, and must offer shared access, like a distributed filesystem.

Once groomed, data is poured into the models, and the training process begins. For the learning process, high-performance access to a variety of concurrent processes is a must, prompting the need for a parallel variant of the distributed filesystem.

Inference cycles are performed on individual nodes drawing on their own dense, speedy storage.

Finding a Method in the MadnessIn a survey conducted by WEKA, a majority of the respondents said that data management is the biggest pain they face for AI and machine learning (ML).

Supermicro and Solidigm have conceptualized a platform that can handle the different AI workloads and meet their changing needs efficiently. It’s a three-tiered arrangement, with GPU-laden servers up top doing the grunt work. These servers make heavy use of High Bandwidth Memory and very fast solid-state drives, courtesy of Solidigm (An example might be the D5-P5430 covered in this article).

The second tier is an all-flash, high-speed storage stack for handling transient working sets for AI and ML solutions. Less compute and inference-intensive than the top tier, these however need much more storage. Once again, Solidigm to the rescue.

ScreenshotThe bottom-most tier is a data lake. This serves as the repository for the source data, if it isn’t culled from other sources (like sensor data, not something scraped off of the Internet). Archived and intermediate results that aren’t being crunched also end up here. This tier may not necessarily be all-flash, but it uses some flash drives for caching.

Making the Dirty Work Clean and GreenHigh on the agenda of Supermicro and Solidigm, is to make AI datacenters green. Solidigm SSDs are already known for their industry-leading physical and power density. These solutions shrink physical footprint by almost 90%, and reduce power consumption by 77%, says Solidigm.

Supermicro is also building systems with incrementally more cores, more PCIe channels, and higher EDSFF drive slots. Additionally, they are running an industry-wide green computing adoption which they predict might save $10B in energy costs each year.

Supermicro platforms now offer GPUDirect storage in which GPUs have DMA capabilities with storage. Combined with higher efficiency processors, these are making power-hungry AI workloads relatively affordable.

ScreenshotSupermicro projects its turnover to reach $14B in 2024, which is double the $7.1B recorded in 2023. In the talk, they attribute this growth to the rising need for AI and ML engines.

ConclusionAI and ML introduces a new and steep set of requirements around processing, memory and storage. Supermicro and Solidigm are making good on their efforts to provide compact, power-efficient, and environmentally cleaner solutions that meet those needs squarely, making infrastructures ready for the future.

Supermicro and Solidigm have a lot more to say on this subject than is covered in this article. So be sure to check out Supermicro’s complete set of AI storage solutions, and for technical details, read their whitepaper on Accelerating AI Data Pipelines. Also check out Solidigm’s presentation on AI storage from the AI Field Day event.


© Gestalt IT, LLC for Gestalt IT: Solidigm and Supermicro Put “Green” in Greenfield

View Details

The rapidly-evolving field of AI relies on powerful computing for tasks like training and making predictions. Up until now, Graphics Processing Units (GPUs) have been the go-to hardware for these workloads. They are powerful, robust and capable of handling complex calculations. But for a minute, let’s explore a viable alternative – the Central Processing Unit (CPU).

CPUs offer significant advantages in cost and availability compared to GPUs. And some new generations of processors packs enhanced compute power to process certain AI applications reliably. But as with most things, there is no one-size-fits-all solution.

ScreenshotDon’t Rule Out the Humble CPU Training AI models typically involves massive datasets and complex calculations. The muscle behind the models, GPUs, have hefty computational and parallel processing prowess to handle this. But now CPUs too are rising through the ranks. It turns out that the modern crop of CPUs proves fairly reliable for a process called inferencing.

Inference deals with using trained models on new and smaller sets of data. The overall computational load is far lower than the training process. CPUs, being more affordable and ubiquitous than GPUs, can be thrown at these smaller problems, especially where the models have been optimized for inferencing on CPU accelerators.

Making the Right Pick is KeyThe choice between a CPU and a GPU for AI inference comes down to three key factors:

BudgetAcquisition costs for CPUs is significantly cheaper taking into account that companies already use CPU accelerators in other aspects of their applications, and they can perform both general compute and inferencing of smaller models or applications with infrequent tasks on them. In the long run, leveraging a single accelerator for varied workloads unlock significant cost-savings.

Workload SizeWhen evaluating workload size, the complexity and size of the model are key considerations. Smaller models have fewer parameters which CPUs can handle efficiently. North of 14 billion, and reaching the trillion parameter range, a GPU is a more suitable alternative.

The frequency of inference is also a deciding factor. For one-time tasks or something that occurs less frequently, CPUs are adequate. But for continuous inference cycles, and applications processing a high volume of users or data streams, GPU’s ability to handle multiple tasks simultaneously is critical.

LatencyRelated to workload size, but more specifically aligned with the application requirements, is latency. CPUs can sufficiently support applications that can tolerate slightly slower processing (inference time), and have proven, for certain workloads, to exceed latency requirements for common user experience guidelines. However, GPU is your best bet if the applications have very low latency thresholds.

The takeaway here is, CPUs are cheaper and more than suitable for trained AI models, especially lighter ones requiring less frequent tasks. By comparison, GPUs are faster, and therefore better suited for highly complex models, big data loads, and situations requiring ultra-fast turnaround.

Intel Is Committed to AIA longstanding leader in accelerators, Intel has big AI ambitions. The company has recently launched its 5th generation of Xeon processors. This isn’t a minor update. Xeon 5th generation reflects a strategic focus that goes much wider than just the generational gains on the silicon.

The 5th gen Xeon is purpose-built for AI workloads from the edge to the datacentre. It delivers a massive 36% performance improvement for specialized AI workloads, unlocked by the Advanced Matrix Extensions (AMX). AMX handles the matrix math operations essential for training and inference. Already proven with large language models of up to 20 billion parameters, it optimizes and accelerates deep learning, delivering impressive vision model performance.

Intel’s commitment to the broader AI ecosystem is equally noteworthy. The company actively collaborates with a growing pool of software vendors and projects to ensure that developers can easily harness the power of the Intel AMX instructions. Their open-source OpenVINO toolkit simplifies AI deployment by optimizing models from popular frameworks like TensorFlow and PyTorch, for diverse hardware platforms.

By actively supporting developers and prepping software for hardware, Intel demonstrates a comprehensive AI strategy that builds upon the strengths of their 4th Gen Xeon processors (Sapphire Rapids) to deliver a powerful silicon-to-software solution for future AI.

Wrapping UpWhile GPUs remain the holy grail for complex AI training, CPUs offer a compelling alternative for the lighter inference tasks. Their affordability, efficiency, and advancements as seen in Intel’s 5th generation Xeon with AMX instructions, make them a strong contender for AI workloads. But the choice between the two boils down to the users’ specific needs, budget, workload complexity, and latency requirements.

For more, be sure to check out Intel’s presentations with partners at the recent AI Field Day event.


© Gestalt IT, LLC for Gestalt IT: Intel Xeon CPUs – How Efficiency Makes It a Top Contender for AI Inference

View Details

Platform9’s Elastic Machine Pool (EMP), is aimed at enhancing the efficiency and cost-effectiveness of Kubernetes operations on AWS. The solution addresses a critical issue faced by many organizations: the underutilization of EKS clusters, which often leads to inflated cloud expenses without corresponding benefits in performance or scalability.

The Challenge of UnderutilizationIn my experience, managing cloud resources efficiently has always been a balancing act. The conventional approach often leads to over-provisioning to ensure availability, inadvertently resulting in low utilization rates. It’s not uncommon to see EKS clusters operating at around 30% utilization, a clear indication of wasted resources and unnecessary costs.

Introducing Elastic Machine PoolEMP represents how resources are allocated and managed in the cloud. By leveraging advanced virtualization techniques and dynamic resource rebalancing, EMP ensures that applications receive the resources they need without the excess that typically accompanies traditional provisioning methods. This approach not only optimizes resource utilization but also minimizes disruptions to running applications, ensuring smooth and efficient operations.

Revolutionizing FinOps and Cloud Cost Management with EMPIn the context of FinOps, Platform9’s EMP stands out as a pioneering tool that bridges the gap between financial management and technical optimization in the cloud. It addresses key issues such as resource fragmentation, sizing mismatches, and the lack of automated feedback loops in traditional EKS setups. By providing a more granular and flexible approach to resource allocation, EMP enables FinOps teams to achieve a finer balance between cost, performance, and availability. This not only leads to direct cost savings but also fosters a culture of innovation and continuous improvement within organizations, as teams are no longer constrained by rigid resource allocations or fear of service disruptions. EMP’s integration into EKS environments exemplifies the synergy between technical innovation and financial stewardship, marking a significant leap forward in the pursuit of cloud excellence.

Cost-Effective Compute Engine for EKSThe financial implications of adopting EMP are significant. Through its intelligent resource management, EMP has the potential to halve AWS costs for EKS users. This is a game-changer for both DevOps and #FinOps teams, who are constantly seeking ways to optimize cloud expenses without compromising on performance or scalability.

Operational EfficiencyBeyond cost savings, EMP enhances operational efficiency. The dynamic rebalancing of resources, based on real-time demands, means that teams can focus more on innovation and less on the manual intervention typically required to manage cloud resources. This efficiency extends not only to the management of resources but also to the overall performance of applications running on EKS.

ConclusionFor those interested in a deeper dive into EMP and its benefits, I recommend reading the full discussion on Platform9’s blog. This exploration provides valuable insights into how EMP can transform the way we think about and manage cloud resources, offering a path to more sustainable and efficient cloud operations.

You can read more about it directly from Platform9 from their Cloud Field Day presentations or on their dedicated landing page on their website.


© Gestalt IT, LLC for Gestalt IT: FinOps Revolution with Cutting Costs and Enhancing Efficiency with Platform9’s Elastic Machine Pool on AWS

View Details

Big, fast, and efficient storage makes the foundation for powerful and adaptable AI systems. It allows you to train on massive datasets, optimize resource allocation to expensive GPU clusters, and prepare for the future demands of AI.

Why Is Big, Fast and Efficient Important?AI models learn and improve through a process called training. It is just one step in the typical AI workflow. Training involves ingesting and processing massive datasets such as text, images, videos, or other formats relevant to the task.

Big storage capacity ensures that there is enough space to house the burgeoning datasets that are utilized in the AI workflow. With traditional hard disk drives (HDDs) topping out around 30TB, NAND storage devices have emerged as the preferred choice. They are over double in size, and have a significantly smaller physical footprint, making a sizeable difference in how storage systems are scaled within the datacentre.

Performance is key in AI. The faster a storage can access and deliver data, the quicker the AI models can train and make predictions. This efficiency is crucial for real-time applications, such as fraud detection, where time-sensitive decisions need to be made. It is equally important during the training phase where large checkpoint files are stored and retrieved from the disk. Slow data access can cost serious money in this cycle due to lost training time. Technologies like NAND storage have significantly faster data retrieval speeds compared to HDDs.

Efficiency is twofold, as it relates to energy-friendly operations, as well as efficient use of space. HDDs rely on spinning platters and magnetic recording heads. The physical size of these components limits how much data they can store in a given space, and the amount of energy they consume. NAND flash storage don’t have spinning platters which translates to lower power consumption and higher energy-efficiency.

NAND flash memory utilizes flash memory chips that are incredibly small, and allows for a much higher density of data storage per unit volume compared to HDDs. This poses a significant advantage for large AI deployments that require constant data access.

Solidigm, a Leader in NAND SolutionsA relatively young company, Solidigm was established in December 2021. With roots running deep in the world of memory and storage, it emerged from a strategic partnership between two industry giants Intel and SK Hynix. Intel’s NAND and SSD business division brought decades of experience in developing innovative flash memory and solid-state drive technologies to which, SK Hynix, a leading South Korean semiconductor manufacturer, added global reach and tremendous manufacturing prowess.

This fusion created a powerhouse in the data storage landscape. Today, Solidigm is a key player in the evolving data storage landscape, and is extremely well-positioned to develop advanced storage solutions. Already the portfolio boasts a large number of market-leading products and solutions. Supermicro, a leading name in high-performance servers, is a top adopter of its SSDs.

ScreenshotSolidigm’s QLC Portfolio for AI WorkThe Solidigm QLC SSD portfolio is designed to meet the demands of capacity and performance needs of AI workloads, all whilst providing crucial performance in workflows like checkpointing and data ingestion and processing.

In particular, the D5-P5336 is a preferred solution for AI tasks. Available in capacities of up to 61.44TB per drive, it comes in a slim E1.L 9.5mm form factor. To give you a sense of the density that can be achieved in this form factor, a 2U configuration can accommodate a whopping 64 disks. That translates to around 3.93PB of raw capacity. There is no getting near to this amount of capacity in such a small footprint when using HDDs. The ruler saves significant money in expensive data centre real estate and related costs like network ports.

The small footprint, however, does not impact performance at all. The D5-P5336 QLC SSD boasts some serious performance numbers to keep those AI workflows running optimally. It offers up to 7,000 MB/s sequential read and 3,000 MB/s write bandwidth. This translates to better user experience and cost-savings which traditional HDDs can’t compete with.

Wrapping upThe exponential growth of AI workloads demands robust storage solutions that are capable of handling massive datasets, and delivering lightning-fast performance. This demand for infrastructure efficiency will push QLCs to the forefront. Solidigm expects QLCs to account for 30% of the drives shipped in 2024.

As AI applications become more complex, and real-time decision-making becomes prevalent, high-density, high-performance storage like Solidigm’s QLC SSD line will take the center stage. By adopting Solidigm’s advanced storage solutions, businesses can gain competitive advantage and significant cost-savings in the AI race.

Faster training times, bumpless real-time operations, and optimized resource allocation contribute to more powerful and adaptable AI systems, and paves the path for ground-breaking advancements in the fields of healthcare, finance, automotives, and more. Solidigm, with its innovative solutions and commitment to pushing boundaries, is well-positioned to be a leader in this exciting future.

For more, be sure to check out Solidigm’s presentations from the recent AI Field Day event.


© Gestalt IT, LLC for Gestalt IT: Solidigm – A Bigger, Faster, and More Efficient Storage for AI

View Details

Modern software-defined storages are increasingly borrowing from the cloud model to design their architectures. This is not an effort to just mimic the scaling flexibility, to make the scale-out model a cloud-scale or web-scale one. But unmistakable resemblances are also being seen in the technical implementation.

The Emergence of a New Generation of Scale-out Software-Defined StorageScale-out software-defined storage solutions aren’t new. To take advantage of NVMe flash, it’s required that the software is designed for high parallelism and ultra-low latency, from the initiator to the target. Of course, a good hardware architecture, and a high-performance network set up with efficient protocols like RDMA, are important supporting ingredients to this.

As workloads and use cases are evolving, for demanding technologies like AI, the need for a new generation of storage solutions is strongly felt.

On November 29th, 2023, Quantum announced the general availability of a new solution, Myriad. Quantum boasts a comprehensive end-to-end portfolio of software solutions known to unlock the true value of data. A company with a pedigree in storage, Quantum offers solutions ranging from archiving to file and object storage, each targeted at a particular set of uses cases. In all its solutions, data is the chief strategy and center of focus.

What Makes Quantum Myriad a Distinguished Solution?Myriad is a new all-flash storage for files and objects aimed at high-performance, AI workloads and and entertainment use cases. The solution boasts a powerful and flexible scale-out cluster architecture. It relies on a 100G Ethernet fabric that interconnects different type of nodes namely, NVMe storage nodes and system load balancer nodes. A management network connects the storage nodes with the control nodes. The solution is based on standard x86 server for the storage node, and whitebox switches for the others.

ScreenshotThe use of whitebox switches is quite an innovative approach, which some cloud providers have set the trend for. An example is Microsoft Azure’s wide use of SONiC. This allows operators to add some logic in the infrastructure part, and optimize the data-path or the manageability in a highly energy-efficient manner.

There is innovation on the software side as well, starting with a new POSIX filesystem designed to store files as objects in a key-value store. Everything is backed by NVMe disks that are divided by zones – as opposed to the legacy block approach – and grouped by zone sets, a concept typical to many cloud architectures.

ScreenshotThe software uses familiar and proven cloud technologies, like microservices and Kubernetes to provide cluster orchestration, a thing that is becoming de facto in new storage solutions.

There is a focus on manageability and simplicity, not only in providing clear and easy user interfaces, but also enabling near zero-touch storage node configuration for new nodes.

A Powerful Data PlatformFilesystem snapshots and clones are intrinsically supported by the write-on-write nature of Myriad’s key-value store. The architecture also enables inline data services such as deduplication and compression, and metadata tagging to accelerate AI/ML data processing which will be available in a future release. As of today, snapshot immutability is not offered, but Myriad’s architecture has room for it to be added at a later date.

For data protection, there isn’t yet an integrated backup solution or a third-party native integration, but Quantum is working toward using other portfolio products to provide both data backup and replication.

Only NFS is available for front-end protocols,, but others are on the roadmap, likely starting with S3 and a CSI integration with Kubernetes.

It will be interesting see where this architecture fits in – computation storage, AI workload or data processing?

To learn more about Myriad’s architecture and why you should use it, explore Quantum’s product page. Also check out this episode of Gestalt IT Rundown for a first impression of the solution.


© Gestalt IT, LLC for Gestalt IT: Quantum Myriad’s Innovative Cloud-Based Architecture

View Details

With Generative AI showing a year-over-year growth of 500%, and companies like Supermicro shipping out 5,000 integrated racks every month, the need for high-performance storage that can work with both CPUs and GPUs is on the rise.

Storage needs to run efficiently, and at speed, so that GenAI models can run optimally without stopping even for a second. Because the second it goes down, production and money are both lost.

So, let’s talk about what GenAI needs to retrieve and store data, and why is it important that all this is organized and cataloged for the future.

Challenges in AI Storage – The Need for High IOPSMassive datasets and complex computations from multiple users at any given time means computers need high Input/Output Operations Per Second (IOPS). It’s no different from when the typewriter keyboard layouts were formed. There were more efficient layouts than QWERTY, the Dvorak keyboard, for example, but the bottleneck was the hammer keys themselves. The typewriter would jam too much because of information coming in too fast from different directions. Hence, the paper was either blank, or the letters were disorganized.

Imagine if the hammer keys, in this case, IOPS, got quietly stuck during an AI computation? If you are running a larger AI query, you may never get the results you desire. Having to restart queries takes time and leads to low-quality outcomes if the AI cannot research the database for answers.

The I/O Blender EffectThe “IO Blender effect” refers to the complexity of Input/Output operations in storage systems, especially in environments running multiple AI workloads simultaneously. Different data access patterns blend together from the applications that are used. This creates a challenge for storage systems to organize for optimal performance. With millions of tiny I/O reads and writes along with NVMe, storage can get faster than the CPUs that are trying to read/write it. Its the typewriter key problem all over again.

The Significance of Data ManagementOrganizing, storing, and processing are the three key pieces of AI model training. It’s important to have proper data management so when someone queries, the answers they get are sound and reliable.

AI can then turn around and create data models for more accurate responses. The goal is to generate actionable AI insights. These insights help produce faster, and more clear-cut answers to queries.

Other ConsiderationsAt AI Field Day in California, Supermicro’s Wendell Wenjen and Paul McLeod demonstrated the hierarchy of the Gen5 storage platforms that are built for AI-centric software-defined datacenters. They include:

  • Single or Twin HDD for high-capacity storage but low AI functions
  • A Hybrid of SSD and HDD for equal AI and storage performance
  • CXL, TLC, and Solidigm QLC SSDs for highest performance of AI queries in GPU clusters

CPUs and GPUs will likely come under a lot of pressure in each of these cases. But, a hierarchy as above will provide high capacity, and keep costs lowdown. It can be tested and repaired, if needed, while AI models continue to be created and run.

Where Generative AI Goes from Here Depends on…Simply put, if data cannot be referenced properly, or if key data is missed because IOPS are slow, or AI insights are incomplete, it can compromise future data. Imagine putting into a model the information – “2+2=4”, and it reorganizes it to “2+2=tangerine” and never corrects it. Billions of queries will have wrong results.

As companies start to rely more on GenAI, every second information is wrong or unavailable, the end result can be catastrophic. That is why storage that is reliable, efficient, and fast is imperative for Generative AI. It’s critical to have optimized storage with high IOPS and effective data management. Otherwise, there will only be pages of incomplete and incoherent information in the archives.

For more, be sure to check out Supermicro and Solidigm’s presentations from the recent AI Field Day event.


© Gestalt IT, LLC for Gestalt IT: Solidigm and Why Storage Is Critical to the Success of AI

View Details

Network practitioners have long realized the benefits of lab environments. Some use them to learn command syntax and practice configuration speed for certifications. Others build partial replicas of production networks to test out the impact of manual configuration changes and device upgrades.

With the promise to increase scale and frequency, and, consequently, the blast radius of these changes, network automation begs a complete replica a.k.a., a digital twin.

At the recent Networking Field Day event, Forward Networks showcased how their solution creates a digital twin of any network to assist with day-to-day operations.

It’s ComplicatedModern networks are complicated. They consist of myriad devices – routers and switches from one or more vendors, firewalls – possibly from a third vendor – and managed with a skillset different than the first two. Load balancers coming from yet another vendor require additional application knowledge. A healthy dose of cloud to that, three doses (AWS, Azure, and GCP) for good measure, and you have something that resembles a Rube Goldberg machine.

It’s not just the network devices that concoct complexity. The protocols that connect them and the security measures that keep them safe are arguably more difficult to understand. IPv4 and IPv6 are the dominant network layer protocols, while TCP and UDP are primary at transport layer, with QUIC gaining ground. IPSEC and TLS provide encryption, and MPLS, GRE, NVGRE, and VXLAN furnish overlays.

Nothing works if the underlying protocols (BGP, IS-IS, OSPF) are having a bad day, or the supporting systems (DNS and authentication) are down. NAT boxes introduce another level of uncertainty for network traffic.

When the network, cloud, security, or infrastructure teams are to respond to an outage, all this complexity and siloed knowledge amount to finger-pointing and protracted MTTR. Because operating a single production network is difficult enough, and maintaining a second network at the same scale and with all the same feature set and functionality is near-unattainable.

Cloning the Network Forward Networks takes a fresh approach to network modelling. Forward Enterprise generates a digital twin of the network by collecting configuration and state data from all packet-pushing devices on premises and in the cloud. A collector with read-only access performs several show commands, and runs the output through a mathematical model. This model analyzes every possible network behaviour, traces where every packet could ever go, and produces a queryable, vendor-independent data model, complete with visualizations.

Knowledge is No Longer SiloedUsing Forward Networks’ Network Query Engine (NQE), a SQL-like language, every user can access network configuration and state information without needing knowledge of specific network vendor’s command syntax. With access to this tool, the application team can determine if their application is blocked without relying on the firewall resource. The security team can pinpoint the location of the IP or MAC address with a simple NQE query.

Personal Assistant IncludedPacket-pumping devices typically contain data that is of interest to people outside the networking team. But the barrier to entry may still be high for some, despite NQE’s resemblance to SQL making it easy to learn. For non-technical staff, learning NQE for infrequent uses is a bridge too far. Alternatively, they are required to submit a ticket and wait for a response.

Forwards Networks has a way around this. Recently, they added an A.I. assistant to Forward Enterprise that creates NQE queries from natural language leveraging generative A.I. The above IPv6 query was produced from a single prompt – “Write a query to find all devices and interfaces that have an IPv6 address configured on a bridge.” In mere seconds, AI Assist queried the digital twin to find the information requested.

A second feature included in Forward Enterprise is Summary Assist. Summary Assist is the inverse of AI Assist, in that, given an NQE query, it spits out a natural language summary of what that query is doing.

NQE has become a team sport with many areas of an organization creating and maintaining queries. Summary Assist provides someone with a quick understanding of a query developed by another that they are trying to modify or add to it.

ConclusionTimely access to precise network information is critical to successful operations, security, and troubleshooting. Forward Networks’ digital twin fully democratizes network data, accelerating accurate responses to any situation or challenge. Where manual data retrieval methods take weeks and months to produce results, NQE reduces the investigation timespan to hours, even minutes in some cases. The new AI Assist brings it down futher to mere seconds, making operations incredibly speedy and efficient.

Check out Forward Networks’ presentation from the recent Networking Field Day event at the Tech Field Day website.


© Gestalt IT, LLC for Gestalt IT: Improving Network Operations with a Doppelganger from Forward Networks

View Details

Even though I feel like I’m well versed in networking there are times when I feel like I’m clueless. It happens when someone asks me how to find important information in a system that I’m not familiar with. I may have spent a large part of my career working with Cisco routers or Aruba access points but if you ask me to find information on a Juniper switch or a Palo Alto firewall I’m going to be struggling to get around. I may know what I want to find or what I need to accomplish but translating that request into something that is able to be executed on the unfamiliar device is a challenge in and of itself.

One of the promises of AI is that it can do the translation for you. I’ve seen countless examples of LLMs and GPT algorithms providing essays for homework and code that looks like it should pass muster when pasted into an IDE. However, I don’t fully trust the output because I didn’t see the steps to get there. I typed a prompt and something came out the other side. Did the system interpret my prompt correctly? Is there some nuance to the language that it missed? Perhaps it was something that I phrased incorrectly that led to the wrong answer. No matter what it was I feel like obfuscating those intermediate steps is a hazard we need to avoid at all costs.

Fast Forward To the Fun PartsRecently, Forward Networks released a new AI Assist feature for their platform that has me excited. I got a chance to see it in action during a demo at Networking Field Day 34 when Nikhil Handigol showed how it works. If you’re not totally familiar with Forward Networks there’s a lot of other great content in the presentation. The short version is that Forward builds a digital twin of your network that allows you to query the model for details about devices and understand interactions so you have a complete picture of what your systems look like.

The star of Forward’s platform is the Network Query Engine (NQE). This is a powerful way to ask questions to get information from your network. NQE allows you to ask about devices and code levels and paths and other critical information. All the things you want to know right at your fingertips. Well, almost at your fingertips. Because NQE is still a language. It has a syntax and a way to structure your query to make sense to the underlying system. Like the above example, you need to know how to format your request into the method the system uses to extract data. If you know you want to see the code levels on all the routers in a given site you need to know what to ask and how to narrow the query so it gives you exactly what you’re looking for.

Anyone with a background in databases knows how hard the process can be. Anyone that has every tried to learn SQL will tell you it’s easy to get the basics but as you add layers of jointing associate entities and filtering information you add layers that not only create failure points but also impact the way the data is presented. If you don’t craft the query in just the right way you could find yourself missing critical data points or incorporating information that is absolutely useless and just muddies the waters.

That’s where the new AI Assist functionality comes into play. You can ask the Forward Networks platform a plain language query and it will go into an LLM and out comes the query you need to run to find the data. Note that it doesn’t produce the results of the query. Instead, it takes your prompt and outputs the NQE syntax you need to use to find the information. Yes, it adds an extra step compared to something like ChatGPT but it’s a very crucial step for a professional setting.

Because AI Assist shows you what you’re about to type in it gives you a chance for a sanity check. You may not be totally familiar with the NQE syntax but being able to see what you’re looking for helps you understand the various fields that can be searched as well as how the query is put together. Even with rudimentary knowledge of the syntax you should be able to spot if something is out of place. More importantly, seeing the query before you run it means that you’ll be able to learn the language faster and understand why it produced a given output.

Bringing It All TogetherThis is a wonderful way to implement an AI function in a platform that provides information about a network or a collection of systems. Rather than just hoping the software is giving you the right data you can audit the query and make sure it’s returning the right results. You get to ensure that you’re getting high-quality information from your system and even be able to tweak the parameters to gain more insights. It’s the way that all systems should help you get to the data you need. Holding your hand at the start but opening up the full capabilities to those that want to put in the time to learn the power under the hood.

For more information about Forward Networks, NQE, and their new AI Assist feature make sure to check out their website.


© Gestalt IT, LLC for Gestalt IT: Using AI to Enhance Forward Networks NQE

View Details

In this Cloud Field Day Tech Note presented by RackN, Adam Fisher discusses how RackN provides IT Ops with a powerful platform to bring infrastructure provisioning up to speed with the demands of modern enterprise IT. The ability for IT Ops to manage infrastructure with cloud-like efficiency powers innovation for the applications that drive a business. The true value of RackN isn’t just for IT Ops or developers; it’s the snowball effect of improved collaboration between groups that ultimately leads to a better bottom line. Digital Rebar is a key to unlocking that value.


© Gestalt IT, LLC for Gestalt IT: Unlocking Developer Efficiency With Self-Service Dev Portals and RackN

View Details

In this Tech Note article from Pure Accelerate presented by Pure Storage, Stephen Foskett discusses how the introduction of Pure Storage's FlashArray//E family signifies a significant leap forward in their pursuit of capacity space and cost-effectiveness within the storage market. With its solid specifications, broad application support, and unwavering focus on efficiency, FlashArray//E solidifies Pure Storage's position as a leading provider of unified storage solutions. By addressing the diverse needs of organizations seeking high-performance, scalable, and affordable storage options, Pure Storage continues to innovate and reinforce their influence in the industry.


© Gestalt IT, LLC for Gestalt IT: Pure Storage FlashArray//E Targets Bulk Storage with Capacity and Affordability

View Details

In this Cloud Field Day Tech Note presented by RackN, Adam Fisher discusses how RackN empowers IT Ops with the consistency, efficiency, and flexibility required to manage modern data centers. Multi-site management in Digital Rebar provides IT Ops with a centralized catalog for all infrastructure deployments. Infrastructure stacks can be deployed across different environments with operational control and security, all managed from one platform.


© Gestalt IT, LLC for Gestalt IT: Infrastructure Pipelines Become Reality With RackN Digital Rebar

View Details

In this Tech Note presented by Pure Storage, Stephen Foskett discusses how Pure Storage is committed to delivering an integrated and user-friendly storage experience, not just a new storage array, which they showcased at Pure Accelerate. While the company emphasized the flash versus disk argument, conversations with customers revealed that their decision to invest in Pure Storage arrays was driven by the ease of use, Evergreen lifecycle, and data center efficiency benefits rather than flash technology alone. The expansion of the E family allows Pure Storage to access new markets that were previously reliant on disk alone. This development brings customers closer to achieving an all-Pure Storage data center, aligning with their desire for a streamlined and efficient storage solution. Pure Storage's complete integrated storage solution, combined with their dedication to advancing flash technology, sets them apart in a crowded market and positions them as a trusted partner for customers seeking a comprehensive storage experience.


© Gestalt IT, LLC for Gestalt IT: Pure Storage: The Power of Flash and Customer Satisfaction

View Details

In this Cloud Field Day Tech Note presented by RackN, Adam Fisher discusses how RackN changed the infrastructure mindset with Digital Rebar as it enables IT to treat bare metal infrastructure as a disposable commodity. Uptime used to be a metric that proved a server’s usefulness and longevity but as systems run and are patched with layer after layer of updates, sometimes a fresh start is better for overall performance.


© Gestalt IT, LLC for Gestalt IT: RackN Brings Bare Metal Into The Cloud Age

View Details

The importance of testing cannot be understated. Network testing is more than just working bandwidth or certifying network components. Applications play an important role in determining how networks should operate. So too does the role of components integrated with each other at the unit and system level. The sponsor of this episode, Keysight, brings a wealth of knowledge to the discussion borne from years of experience in the networking testing space. Learn how Keysight uses that knowledge and experience to the modern world of high speed Ethernet and the needs of companies deploying it at scale.


© Gestalt IT, LLC for Gestalt IT: Network Testing Is Critical

View Details

In this article presented by Pure Storage, Stephen Foskett discusses the FlashArray R4 platform and how it can empower organizations with high-performance storage arrays tailored to their specific requirements.


© Gestalt IT, LLC for Gestalt IT: Pure Storage Unveils Flash Array R4: Empowering Next-Generation Storage Solutions

View Details

In this Tech Note presented by Solidigm, Max Mortillaro discuses how the Solidigm P5-D5430 QLC SSD drive establishes a new milestone for QLC flash. Its massive improvements in endurance and capacity demonstrate that QLC flash is the new normal in mainstream, read-intensive workloads. Clearly aimed at large datacenters and hyperscalers where storage density, outstanding reliability, and low cost per GB are key factors, the P5-D5430 QLC SSDs is a key enabler to even more robust and massively scalable data stores, allowing better energy efficiency and storage density.


© Gestalt IT, LLC for Gestalt IT: Solidigm’s P5-D5430: QLC Flash for Robust Scalability and Reliability

View Details

In this Cloud Field Day Tech Note presented by RackN, Adam Fisher discusses why Digital Rebar from RackN is such a powerful product. It can be the glue for IT Ops to empower developers to focus on driving business value while providing a consistent infrastructure management platform. In the following posts, I will highlight how the bare metal provisioning, infrastructure pipeline, and self-service portal backend use cases for Digital Rebar prove that RackN has brought infrastructure provisioning to the modern age.


© Gestalt IT, LLC for Gestalt IT: RackN Bridges the Gap Between People and Platforms

View Details

in this Edge Field Day Tech Note presented by Mako Networks, Brian Chambers discusses how segmentation is key to operating at scale and how Mako Networks has made a successful implementation of an edge computing solution. 


© Gestalt IT, LLC for Gestalt IT: Segmentation is a Key Edge Building Block with Mako Networks

View Details

In this Cloud Field Day Tech Note presented by Forward Networks, Justin Warren discusses that in today’s climate, enterprises need to be equipped to handle a diversity of environments without being constantly bamboozled by complexity. An eye to distinguish necessary and valuable variation from the insecure or dangerous ones helps build this ability. With its high-fidelity observability, Forward Networks helps them get the benefits of the new and the old, without making security an impossible task.


© Gestalt IT, LLC for Gestalt IT: Seeing through Hybrid Multi-Cloud with Forward Networks

View Details

In this Storage Field Day Tech Note presented by StorPool, Chris Childerhose discusses how StorPool disrupts the storage scene by unlocking effortless implementation, deployment, maintenance and upgrading, for the first time, all without the slightest service interference to the clients and employees. It’s time to change the way we think about storage!


© Gestalt IT, LLC for Gestalt IT: StorPool – Storage Delivery with Commodity Hardware the Right Way!

View Details

In this Cloud Field Day Tech Note presented by Forward Networks, Remington Loose discusses how Forward Networks helps newly-joined teams work better together, regardless of their chosen platforms or deployments. The reduction in risk to both the overall environment and individual changes can smooth and accelerate the combining of two networks. Given how much technical debt is commonly incurred in these situations, the product can help to avoid the permanent, short-term solution.


© Gestalt IT, LLC for Gestalt IT: Easy, Verified M&A with Forward Networks

View Details

In this Tech Note presented by Portworx by Pure Storage, Joey D'Antoni interviews Venkat Ramakrishnan and discusses why databases are one of the essential parts of an application. But they can also frequently be a performance bottleneck for the entire application, and protecting their data is critical to keep operating. With the increased number of database solutions, and the various deployment and management methods in the picture, both on-premises and in the cloud, it has been challenging for operations teams to stay afloat. Portworx Data Services aims to meet this need by simplifying deployment, reducing administrative toil, and helping applications move forward wherever one wants them to run.


© Gestalt IT, LLC for Gestalt IT: Portworx Data Services Makes MongoDB Simpler

View Details

Advanced SSDs are key in edge environments, offering high performance, reliability, energy efficiency, and storage density. Solidigm's D5-P5430 SSDs, built on reliable QLC flash, provide an efficient solution for maximizing capacity at the edge. These SSDs enable faster speeds, greater endurance, lower power consumption, and compact size, unlocking the full potential of edge computing.


© Gestalt IT, LLC for Gestalt IT: The Value of Cutting-Edge Storage in Edge Environments: Solidigm D5-P5430

View Details

What are the biggest roadblocks to automation in 2023? In this interview with Omar Sultan of Cisco we explore the challenges of automation at all levels of the company as well as the need to focus on the investment and payoff instead of the technical decisions that need to be made.


© Gestalt IT, LLC for Gestalt IT: Making Automation Accessible with Omar Sultan of Cisco

View Details

In this Tech Note presented by Pure Storage from Pure Accelerate, Joey D'Antoni discusses how Pure Storage has used data science and its best-in-class customer support to help users quickly identify and recover from ransomware attacks. Building this intelligence into the Pure ecosystem helps them plan data protection and capacity without complexity. 


© Gestalt IT, LLC for Gestalt IT: New Releases to Optimize Ransomware Resiliency at Pure//Accelerate 2023

View Details

In this Storage Field Day Tech Note presented by StorPool, Denny Cherry discusses how StorPool has brought to the IT enterprise storage space what users have been wanting for  over a decade now - quality, high-speed, highly scalable storage on commodity hardware. With the networking technology having reached the speeds that we need-, and the SSD/NVMe perf at where we need- it, the StorPool software is the last missing link required to manage the environment efficiently and present the storage to the servers using all standard hardware technology.


© Gestalt IT, LLC for Gestalt IT: StorPool Cracks the Commodity Hardware Nut

View Details

In this Storage Field Day Tech Note presented by StorPool, Jim Czuprynski discusses how radically different StorPool's storage model is. By standardizing storage nodes on commodity hardware, and concentrating on provisioning capacity in an as-a-service orientation, it’s much simpler for storage to become and remain what everyone in IT wants - always on, non-disruptive, reliable, and performant.


© Gestalt IT, LLC for Gestalt IT: Getting Off the Storage Refresh Treadmill with StorPool

View Details

In this Networking Field Day Tech Note presented by Catchpoint, Peter Welcher discusses how Catchpoint can monitor and display valuable BGP information over time, allowing staff to keep an eye on, spot problems in, and troubleshoot global or large scale routing.


© Gestalt IT, LLC for Gestalt IT: Catchpoint BGP Monitoring

View Details

In this Cloud Field Day Tech Note presented by Forward Networks, Chris Grundemann discusses how Network observability serves as a linchpin for maintaining a secure and resilient network infrastructure. In the realm of multi-cloud, achieving that observability is even more challenging due to the absence of unified visibility across cloud platforms. Forward Networks has risen to that challenge by extending their digital twin technology to major cloud providers. Their solutions not only deliver deep observability, but also enhance security through tools like Path, Posture, and Blast Radius. With Forward Networks' comprehensive approach to multi-cloud observability and security, businesses can unlock the full potential of their networks, and ensure a robust and protected digital ecosystem.


© Gestalt IT, LLC for Gestalt IT: Multi-Cloud Security Requires Multi-Cloud Observability with Forward Networks

View Details

In this Edge Field Day Plus article presented by Mako Networks, Ben Young discusses how Mako Networks has positioned themselves as a must-have solution in the retail edge space. Having PCI compliance taken care of at the network layer is a massive burden for organizations and often a stumbling block in maintaining compliance. Mako Networks help offload that with ready compliance. Its centralized management reduces infrastructure workloads while saving time with the ability to push out templated policy changes and software updates through a single portal.


© Gestalt IT, LLC for Gestalt IT: Mako Networks – A World-Class PCI-Certified Network-as-a-Service Ecosystem

View Details

In this Edge Field Day Tech Note article presented by Mako Networks, Carl Fugate discusses how professionals working with SD-WAN and SASE offerings on a daily basis will find Mako Networks' almost obsessive focus on differentiating their products impressive. They have uniquely positioned themselves as a must-have solution for retail environments and sectors where governance and compliance are critical requirements. The multi-operator management model which enables independent management of different network zones at a site is a feature not encountered on other like products. Hopefully, more companies will follow Mako’s lead in targeting features that support specialized use cases in the future.


© Gestalt IT, LLC for Gestalt IT: Mako Networks – Security-First Networking at the Edge

View Details

I recently wrote about Catchpoint’s presentation at Network Field Day 29, That blog compared Catchpoint to some competing products or at least other products that include some similar capabilities. Catchpoint’s stated goal is to be the Best at Internet Resilience.

This article provides a mildly deeper dive into products’ relative capabilities, and then looks at Catchpoint from the Internet Monitoring and Resilience perspective, trying to answer the question: What are the top differentiating capabilities Catchpoint provides for Internet Monitoring, etc.?

We’ll finish up with a Troubleshooting Use Case walk-through, looking at how the readily available reports in Catchpoint solved an actual SaaS app slowness problem.

Network/App Response Monitoring ProductsIn the network and application performance monitoring (etc.) space there are several products. There are overlaps in some or many capabilities, but it can be hard to tell the products’ capabilities apart. However, Catchpoint differs in that it’s a solution that monitors your entire internet stack, rather than just applications and network traffic. To that end, it’s focused on the experience of the user through the entire digital service delivery chain.

Take a Closer Look with Catchpoint: Synthetic Monitoring Live Demo – Networking Field Day 29Catchpoint’s PositioningCatchpoint’s intent is to be the best at Internet Resilience (including monitoring and reporting) overall.

That’s quite a goal! The word “resilience” suggests “fast troubleshooting” to me, but in this context it’s about developing a comprehensive monitoring strategy that enables predictive insights, contingency planning and continuous improvement over time. A company with a resilient internet should be able to proactively deal with issues before they impact their users and be able to implement alternative paths that prevent outages. We’ll take a look at how they do that below.

Why Best at Internet resilience? Well, just about everything lately depends on the Internet. Especially delivery of Internet application content to customers and WFH staff. Or Internet access to office apps from home or while travelling.

Datacenter, CoLo, and Cloud traffic between apps and data storage may travel over dedicated links. But dedicated links are subject to much less variability. So, the entire Internet is the bigger challenge, and requires an application focus. Catchpoint of course handles other types of links as well.

What are the Top Catchpoint Internet Capabilities?I asked Catchpoint about this: what are the top Catchpoint capabilities regarding Internet Resilience?

The answer:

  • Catchpoint has the most global monitoring points for customers to leverage. At the time of this writing, (2022) Catchpoint has over 2000 vantage points across the Internet, including monitoring points in China which makes them less impacted by the Great Firewall.
  • Catchpoint can do outside-inwards monitoring/troubleshooting, e.g., from the desktop of a user travelling in Dubai having problems accessing corporate apps or SaaS. This can provide data on problems invisible to competing products.
  • Catchpoint provides focus: monitoring, reporting, and alerting on issues you care about. E.g., the latency of an application to, say, 3 locations. It does RUM (Real User transaction Monitoring if desired. It also does synthetics. I’ve been cautioned that “synthetics” means different things to different vendors. Catchpoint can do basic RUM (web URL request components: DNS, connect, SSL establishment, wait time, etc.). Or “synthetics 360”: deeper dive diagnostics, providing developer support. (For more on this, see also this set of web pages.) Catchpoint monitors key statistics Google uses to score websites. And Catchpoint has plug-ins for the Chrome and Edge browsers that can capture additional useful data.
  • Catchpoint reports “User Sentiment” – monitoring third parties for user comments, down detection, etc. In other words, reporting on external perceptions of performance as well as the hard network performance data points.
  • Catchpoint includes an Internet Weather Report and can monitor as much or as little of the Internet as the customer wishes (and wishes to pay for).
  • Catchpoint is independent, not tied to a hardware vendor or other functions. It is agnostic about your network, server, cloud, CoLo, and other brands.
  • Expertise. Catchpoint includes consulting services by a value engineer and a customer service engineer to new customers. This ensures identifying key use cases, setting up data collection, and demonstration of how to use the various relevant reports Staff augmentation or further consulting services are also available.

What else makes Catchpoint different aside from its enormous global observability network? One item mentioned is finer data granularity, with no limit on retention.

Note also that Catchpoint has a lot of basic to intermediate documentation about various skills relating to monitoring and troubleshooting. This includes some good how-to documents.

About Catchpoint’s MeasurementsCatchpoint’s “web RUM” allows monitoring of the various components of a web-based application. Consider that such an application might have multiple global users accessing micro-services scattered around various sites.

Public-facing web apps require many Internet services. That starts with DNS, but includes Content Delivery Networks (CDN’s) as well. If their response is slow, the application experience will be degraded. One common potential problem is mis-configured global CDN services, where a user is hitting a CDN in remote location, rather than a closer one.

Another key set of measurements relate to Internet path. What path is a given user’s or site’s traffic taking, and is some segment of that path performing poorly?

In relation to Internet path, BGP is of course of major interest. Monitoring BGP peering, Internet paths for various IP prefixes, etc. is important. Catchpoint’s capabilities around BGP and Internet paths will be discussed in a follow-on article.

Catchpoint does also provide end-user WiFi and other monitoring for user-centric problems. But that’s a topic for another time and article.

The reason we use networks is of course applications, in the broad sense. Catchpoint can monitor:

  • DNS
  • CDN’s
  • BGP
  • Voice (VoIP) and video quality
  • SaaS apps (like Salesforce)
  • Cloud apps (like Teams and Zoom)
  • API’s

To sum up, Catchpoint monitors web application and component performance.

One valuable use of Catchpoint is working with the application owners to determine the slowest components to load, with the objective of speeding up page loads.

The following example walks through troubleshooting an actual problem. One consequence of finding the cause might be a change in the web page to improve performance under adverse conditions.

Take a Closer Look with Catchpoint: BGP Monitoring Live Demo – Networking Field Day 29Internet Troubleshooting ExampleI’m a classic techie: show me some details! I hope you feel the same way. Let’s look at some detail!

Catchpoint walked me through a real-world example. They have a SaaS CRM app that we won’t name used by some of their staff. (Note: all of their staff is WFH, Work From Home or wherever.)

The app had some problems. This section showcases the data provided by Catchpoint. We start with an overview page.

Notice the first “tile”, top-left in the following screen capture (I added the red rectangular border).

Catchpoint measures the availability of applications from different vantage points. It provides insight directly on the backbone of Internet from Data Centers and ISPs all over the Globe. Many times SaaS applications are available from the vendor’s perspective but users are still unable to access them. This gauge represents the reachability to actual end user devices regardless of their location.

Looking at the two bottom left tiles (see below for close-ups), red is bad, as usual: higher page load times and page load failures.

The bottom left tile shows synthetic monitoring from the end user device, the next tile to the right shows actual real user experience as rendered in the browser. If you home in on the time (x-axis), the left tile shows problem onset before the middle one: apparently there were no real users of the app at 9 AM.

Then there’s the right bottom tile:

The bottom right tile shows synthetic testing response times as measured from the Enterprise Nodes. These would be agents on nodes in the organization’s network, but dedicated machines with no user workload. They can do browser emulation.

So why are the red datapoints on the right lower? Answer: failure is faster. Those web page loads failed to complete and timed out.

The upper world chart tile shows real user performance, color-coded.

This helps you quickly see where affected real users are.

Hovering over a data point, brings up some additional data. (Not shown here.)

But there’s even more data readily available!

Clicking on a data point and then on the 3rd icon on the left side bar brings up a different set of information, shown below.

The top shows a timeline of measurements. Selecting one shows further data below the timeline: What the web page looked like, and various statistics. The bottom “filmstrip” shows what the web page looked like at various times as it rendered. This can help you see where things slowed down or failed.

If you look again at the top set of data points over time, notice that some successful page loads were twice as fast as others. This suggests some random delay somewhere.

Scrolling down provides more information, a waterfall chart of objects loaded. The highlighted line shows a moderately long wait for that object.

Clicking on that brings up more details:

Scrolling further down, we see:

The video in row 32 is taking about 4 seconds to load. Slow!

And that was the actual problem: the web page includes a MP4 video clip. It is larger content and was taking a longer than usual time to download and render.

This provides the basis for an informed decision. You (and any other parties involved) might consider:

  • Do you put up with occasional page fails (timeouts) due to the video clip when some sites or regions are experiencing congestion and slowness?
  • Or do you remove the video clip, or reduce its resolution to reduce the number of bytes transferred, etc.?

That also gives the flavor of how this might be used in developing or maintaining a web app: use the RUM data to optimize web page load times, etc. Additionally, it allows you to hold SaaS Vendors accountable for content on their sites. Catchpoint RUM data shows how many times a page was loaded over time so quantifying the impact that a slow page has on employee productivity is easy!

If you go back to the original screen and then scroll down, there was more data available there.

The upper right tile clearly shows a spike in average load times.

The bottom left tile breaks out load times by ISP, with the small squares representing one user. This helps spot where there is an ISP problem. The bottom right tile summarizes the per-ISP data in text form.

Catchpoint’s Day 1 SupportI’m seeing a growing trend in networking and IT awareness that buying software or doing automation can result in “shelfware”. Staff needs to know how to use the new tool and have a good reason to do so.

As part of its process, Catchpoint includes services by a “value engineer” as well as a “customer success engineer”. The value engineer’s job is to understand top use cases and get the customer set up to monitor what is needed and understand the reporting.

Deeper dive and site staffing are also available for a fee.

ConclusionCatchpoint provides a lot of reporting out of the box. I’ve found myself somewhat overwhelmed by fast presentations (e.g., at Network Field Day). Catchpoint staff walking me through the above use case helped, and I’ve tried to share that in this article.

My hope is that sharing that experience with you, the reader, will help you envision how Catchpoint might be useful for monitoring your Internet stack as well as other applications, re-engineering them to be more tolerant of slowdowns, and respond to outages more rapidly through well-organized pre-analyzed data. Catchpoint makes your Internet more resilient which requires more than traditional application or network performance monitoring. The breadth and depth of monitoring capabilities shown to me suggest that they can do just that.

Watch all of Catchpoint’s videos and demos from Networking Field Day 29 on the Tech Field Day website.


© Gestalt IT, LLC for Gestalt IT: Catchpoint Excels at Internet Resilience

View Details

At the recent Storage Field Day event, Solidigm presented the new D5-P5430, a massively dense NVMe QLC SSD available on the new E3.S EDSFF, U.2 and E1.S form factors. The launch was coordinated with Supermicro, an early adopter of the EDSFF specification that recently introduced a new server family offering support for E3.S SSDs.

Top FeaturesFrom a capacity standpoint, the solution is very appealing. The new range of Supermicro servers accept up to 30.72E3.S EDSFF SSDs. Considering the announced capacity of 30.72 TB per single Solidigm D5-P5430 E3.S SSD, a fully populated 2U chassis provides 983 TB of raw capacity.

The scalability benefits of the solution are noteworthy too. The EDSFF form factor was developed by an industry consortium in collaboration with the Storage Networking Industry Association (SNIA) to address longstanding issues and inefficiencies of standard storage form factors that were inherited from the mechanic hard disk drive era.

What’s most interesting about the density achieved by these next-generation architectures is the set of sustainability benefits that comes out of it. The recent Storage Field Day 24 industry event sparked some very interesting conversations on the topic.

Three Goals of SustainabilityThe impact of recent energy price hikes is felt across organizations, especially with those running infrastructure at scale. To counter that, IT teams have started scrutinizing energy efficiency of all infrastructure components, starting with CPUs and storage.

From a practical standpoint, organizations look at achieving three sustainability objectives:

  • cut back energy consumption
  • use greener energy sources
  • reduce greenhouse gas emissions (GHG) as a compound of the first two objectives

In order to achieve a low energy footprint, they must understand the energy efficiency of a solution and its overall impact when operating at scale. Putting performance and impact of load levels aside, several factors come into consideration when looking at the basic building brick of a storage infrastructure (in this case, Supermicro’s 2U chassis populated with Solidigm D5-P5430 E3.S SSDs).

The capacity and the physical footprint allow us to determine the density. The more capacity can be packed into a smaller footprint, the denser a storage system is.

Storage density is important from a sustainability perspective. Where capacity is the deciding factor over performance, the Watt/TB measurement helps determine how energy efficient a given storage system is.

At similar price points, using a single 2U server to achieve 960 TB raw capacity with an arbitrary power draw of 1000W (taking a value out of the blue) is way more sustainable than using two 1U servers with 480 TB raw capacity each, and a power draw of 700W each (again, random value).

Besides power draw, thermal dissipation requirements grow with the number of servers deployed. Thermal dissipation primarily depends on datacenter HVAC systems, and therefore snowballs energy consumption.

While this may not be important for single digit deployments, it becomes a critical issue when operating at multi-petabyte scale. Increasing the storage density allows organizations to fulfill capacity requirements with less hardware, and thus reduce thermal and energy requirements compared to non-optimized form factors or previous-generation hardware.

Improved energy efficiency will also lead to a reduction of GHG emissions and help improve ESG reporting. Combined with other measures (for example using renewable energy sources to power their data center), this can lead to more sustainable operations.

Additionally, there is the matter of future scalability. A 30.72 TB E3.S SSD may be impressive today, but in a couple years the capacity of E.3 SSDs could significantly increase. This will lead to even denser systems, greater energy efficiencies, and further GHG emission reductions.

ConclusionThe combination proposed by Supermicro and Solidigm enables a seamless multi-petabyte scale storage that is ideal for organizations like hyperscalers, xSPs, and large enterprises with an ever-growing need for capacity-oriented storage. Alternatively, from a sustainability perspective, it ensures greener storage deployments with a significantly reduced energy footprint.

To listen to the conversation, check out the Gestalt IT Roundtable featuring Supermicro and Solidigm, recorded at the recent Storage Field Day event.

For a better understanding of the scope of sustainability in enterprise storage, be sure to check out “Measuring Sustainability in Enterprise Storage” a story that discusses this in detail.

Watch the full discussion on our website with Solidigm and Supermicro.


© Gestalt IT, LLC for Gestalt IT: Solidigm and Supermicro Help Organizations Achieve Three Goals of Infrastructure Sustainability

View Details

As software introduces better consumption models for underlying technologies, the appetite to consume more resources continues to grow. Composable infrastructure feeds this hunger by serving up resources as and when they are needed.

At the recent Storage Field Day event where Solidigm presented its QLC SSDs, composable architecture took centerstage.

A veteran of the SSD industry, Solidigm recognizes that CPU and memory have largely outpaced storage in terms of speed and density increases. Newer configurations are working on larger datasets, and many of the workloads – AI and machine learning – are data-centric. Their progress is limited by their ability to get to the data.

Solidigm’s use of Enterprise Datacenter Standard Form Factor (EDSFF) truly stands out in this context. EDSFF is front-loading and hot-pluggable, and provides better density options than the standard drive form factor.

Solidigm D5-P5430At the post-event roundtable hosted by Solidigm and Supermicro to kick off the Solidigm D5-P5430 offering, Soildigm’s Tahmid Rahman and Supermicro’s Patrick Chiu fielded questions about the product and Supermicro platforms designed to work with its high-density drives.

In a nutshell, the D5-P5430 is available in a U.2, E1.S and E3.S form factor and in a range capacities from 3.84-30.72TB IT is a storage device that communicates via NVMe over PCIe gen 4.0. It’s one of the first devices to be offered in the EDSFF E3.S package.

The E3.S form factor is designed for both density and efficiency. Tahmid and Patrick claim five times the density of traditional SSD packaging, and almost twice the power density with it.

The form factor is especially valuable for composable infrastructure, due to being hot-pluggable and front-loading which allows users to get the right size for their needs by just adding and removing drives.

New Platforms from SupermicroSupermicro has designed a set of new chassis to accommodate the drives, allowing incredible density. Their offerings provide 16 E3.S slots in a 1U chassis, or 32 in a 2U configuration. (Real estate on 2U puts the absolute maximum count at 40).

The newest platforms are designed around a single AMD Genoa processor, and alternatively, a pair of Intel Sapphire Rapids CPUs. The AMD Genoa sports 128 PCIe 5.0 lanes, while the Sapphire Rapids have 80 lanes each. That’s more than enough to handle the full bandwidth of the maximum number of drives available in the chassis, with plenty to spare for network connectivity.

These numbers help explain why the density available on this platform with the Solidigm SSDs is useful. This much density is essential to keep pace with the throughput available on the newest generation datacenter processors.

Tahmid points out that over recent history, storage expenses are outpacing compute by 4.5x. These drives are helping storage density stay in the race with CPU and memory advances.

Real-World UsersWho has need for this type of composable infrastructure? That would be top-tier enterprises, when dealing particularly with large datasets. Hyperscalers, content-delivery networks and streaming media services gobble up storage prodigiously. These systems will be ideal for their scale and use cases.

It’s also an ideal workload for QLC technology as these are largely read workloads. Tahmid emphasizes that on a realistic content delivery network workload these SSDs can be written at a rate of 2.4TB per hour for the full 5 years without wearing them out.

Solidigm also highlighted a study, driven by the University of Toronto where they found nearly 99% of SSDs taken out of service were at less than 15% of their wear limit.

ConclusionIn modern IT, one size doesn’t fill all. Sometimes the users need more. Designed to deliver more, composable architecture allows for the existing pieces to be fit together in new ways that expand the capabilities of systems. As a concept, composable architecture has come of age.The Solidigm D5-P5430 is proof of that. With its blend of performance, density, affordability and optimized form factor, it has the potential to become a key component in right-sizing datacenter infrastructure.

Find out more about the Solidigm Data Center SSD offerings on their site or contact them for more information. You can also watch the Gestalt IT Roundtable Discussion available now.

Another resource to check out in this connection is – CXL meshes nicely with Composable Infrastructure.


© Gestalt IT, LLC for Gestalt IT: Composing a Harmonized Infrastructure with Solidigm and Supermicro

View Details

As enterprise and cloud data growth continues, opportunities are appearing to leverage next-generation SSDs in new server form factors to improve capability and sustainability. This Gestalt IT Tech Talks Roundtable features Tahmid Rahman of Solidigm, Patrick Chiu of Supermicro, discussing next-generation SSD with Max Mortillaro, Andy Banta, and Stephen Foskett of Gestalt IT. Today’s SSDs have a new form factor called EDSFF or E3 for increased density as well as next-generation NVMe protocol support for high performance, low latency, and new capabilities. Solidigm is developing specialized SSDs to meet real-world workload demands from AI and analytics applications with the right mix of performance and density. Specific use cases include dense object storage, content delivery networks and streaming media, and online application processing. D5-P5430 is the first E3 SSD from Solidigm with improved density, thermal efficiency, and signal integrity for future PCIe Gen 5 and 6 products. This includes the use of PCIe Gen 4 and QLC NAND flash for increased density. The newest server platforms from AMD and Intel include more PCIe lanes, enabling Supermicro to make the most of this new SSD technology in ultra-dense high-performance servers.

Panelists for this Solidigm Roundtable:PanelistsMax Mortillaro

Andy Banta

Twitter@MaxMortillaro

@AndyBanta

Solidigm PanelistTahmid Rahman, Director, Product Marketing, Data Center Group at Solidigm. You can connect with Tahmid on LinkedIn and find out more about Solidigm’s products on their website.

Supermicro PanelistPatrick Chiu, Technical Director – Product Management, Business Development and Strategy at Supermicro. You can connect with Patrick on LinkedIn and find out more about Supermicro on their website.

ModeratorStephen Foskett

Twitter@SFoskett

Key PointsIn this Gestalt IT Tech Talks Roundtable, Tahmid Rahman of Solidigm and Patrick Chiu of Supermicro join Max Mortillaro, Andy Banta, and Stephen Foskett to discuss the opportunities presented by next-generation SSDs and new server form factors. The focus is on leveraging these advancements to improve capability and sustainability as enterprise and cloud data growth continues. Solidigm is actively developing specialized SSDs, such as the D5-P5430, to meet real-world workload demands, particularly in AI and analytics applications. These SSDs, utilizing the EDSFF (E3) form factor, offer increased density and support next-generation NVMe protocol for high performance and low latency. Supermicro, taking advantage of the latest server platforms from AMD and Intel with more PCIe lanes, aims to maximize the potential of this new SSD technology in ultra-dense, high-performance servers.

The discussion highlights key challenges faced by the industry, including data growth, hardware modernization, scalability, performance, efficiency, and cost optimization. Solidigm and Supermicro are actively developing advanced storage solutions, including the next generation of SSDs and Senhupascale storage, to address these challenges. The new storage solution being developed utilizes various form factors such as U.2, E1.S, and E3.S, offering unique features and innovations. The focus is on the EDSFF form factor, which is highly adaptable for high-frequency interfaces and offers improved thermal efficiency, better density, and scalability, resulting in reduced power consumption and environmental benefits. The combination of this new form factor and Supermicro’s expertise enables applications like dense object storage, content delivery networks, and streaming media.

The introduction of the EDSFF form factor aims to increase storage density and reduce power consumption, aligning with the goals of optimizing IT infrastructure, reducing operational costs, and meeting the growing demands of data centers. The new SSDs utilizing EDSFF offer higher power density, improved energy efficiency, and align with sustainability objectives. The market is expected to gradually adopt EDSFF, with hyper-scalers and storage unit vendors leading the way. By 2026, EDSFF adoption is predicted to exceed 50% of the market, providing significant performance and density benefits for applications such as content delivery networks, object storage, AI, and general-purpose servers. Solidigm, in collaboration with Supermicro, is excited about the potential of the EDSFF form factor for storage, as it offers enhanced performance, density, power efficiency, and high-performance AI interfaces. The adoption of EDSFF by customers will be driven by the compelling advantages provided by complete system solutions.


© Gestalt IT, LLC for Gestalt IT: Bringing Next-Generation SSD to Enterprise and Cloud with Solidigm and Supermicro

View Details

With news of ransomware attacks forcing schools, hospitals and large corporations offline for days, organizations are anxiously counting days to their predicament. Anybody who has followed the chronicles of cyberattacks over the last few years know that it’s a matter of when, not if a company’s data is locked up in a ransomware prison. Making things worse is the uncertainty of whether its a ransomware incident on company systems, client systems, or those of the vendors that is going to bring a business down.

Traditionally, enterprises have used anti-virus and malware detection to protect against attacks, but the speed and methods of ransomware have evolved over the years. Legacy methods don’t recognize zero-day techniques, and are often slow to detect well-crafted cyber weapons.

All we know about these ransomware tactics, techniques, and procedures (TTPs) is that all parts of their systems have become faster, especially encryption. The fastest TTPs encrypt only a portion of the file, and use techniques other than standard AES stream ciphers, to ensure they can make the largest number of files unavailable the fastest. In recent tests, over 200,000 files could be encrypted in four and a half minutes.

Insider Threats, Both Malicious and StupidSome internal data breaches happen without malicious actors. Once, a co-worker who could not work with data in its system of record, started downloading and exporting it into spreadsheets where he created all kinds of amazing formulates, pivot tables, and linked tabs. These spreadsheets, stored on his local desktop, were not part of any regular backup system. If he wanted to work outside the office, he would copy these spreadsheets to a USB drive to take them home. He occasionally lost thumb drives and told no one. The only way the company found out he was doing this was because he kept asking for replacement USB drives.

At another client, a contractor stole hardware from the company due to sloppy asset management processes. To make matters worse, he was told that he could make much more money selling customer data from a national membership-based retailer to which most people had given sensitive data. Unfortunately, when he was caught, it wasn’t for the data, but by the local police, who raided his home for other thefts. No one at the company knew data had been copied.

RackTop Systems BrickStor SPI recently attended a briefing session with RackTop Systems on their BrickStor Security Platform (SP) for unstructured data. In the session, they demoed running a live ransomware attack against a test system (YouTube link to a similar demo). They were able to shut down the malware, block access to the logged-in user who was the entry point for it, and restore the previous versions of the files, all within minutes. The only reason it took minutes, not seconds is that they wanted to show us the impact of the ransomware and the actions BrickStor SP took as it was doing its job.

How did it do this? BrickStor SP evaluates user behavior, account permissions, file activity, and other metadata about the files and the user in real time. They call this approach the “Data-Centric Zero Trust Policy.”

The differentiator is the analysis of user behavior. Not just data about access permissions, BrickStor analyzes whether entities are accessing files they don’t normally access, the sort of actions they are taking, and if it’s normal.

It detects user behavior anomalies via User and Entity Behavior Analysis (UEBA). Once detected, the system can take over, stopping all access by that user that was the weak point. A security operator can review and approve the next steps because the attack has been eradicated via that vector.

Why Is Active Defense Important? Traditional detection takes longer, and stopping access isn’t always automatic. Throw in the fact that “just restore” approaches can take days; the encryption needs to be stopped before petabytes of data are restored. There also has to be protection against the initial data exfiltration that comes with double extorsion approaches now being rolled into ransomware incidents.

Preventing Data ExfiltrationCircling back on the real-life stories from before, BrickStor SP can stop careless employees from downloading files or files at the volume they shouldn’t be downloading. Cherry on top, it monitors user behavior, stopping an insider from downloading files and putting them somewhere they should not be, including a USB drive, just because they like working on unsecured devices. It can also stop malicious insiders from exfiltrating large volumes of data. The key is that it isn’t doing this based on some set of general policies but on actions by a particular user and a specific file set.

Summing UpThis article touches only a few of the features that BrickStor SP offers in interest of time and space, but here’s what I think. The greatest merits of BrickStor SP are active, data-focused detection based on advanced analytical approaches to user behavior, speedy detection and lock-out that saves time and money for the resumption of normal operations, and integration with AD and LDAP. The product covers both insider and external threats, malicious and ill-thought-out actions, and most importantly combine storage and threat detection.

The only downside is that BrickStor SP protects unstructured data for SMB and NFS file shares only, These are drawbacks only to someone searching for solutions beyond storage. They only offer support for data classification from third-party solutions, but that is a common way to address data categorization and classification and most organizations already have solutions in this area.

If you’re shopping for a cyber storage system for unstructured data, take advantage of the virtual BrickStor SP free trial on their website.

Watch the RackTop BrickStor Showcase and more on the Tech Field Day website for more information.


© Gestalt IT, LLC for Gestalt IT: RackTop BrickStor – Protecting Data Against Ransomware

View Details

If I asked you to build a table you might be able to figure it out. You’d look for drawings or some other kind of plans online and get to work. Maybe you’re the kind of person that will just start cutting some lumber and figuring out how to put it together. If I asked you to build a computer desk with integrated power, motors to raise it to a standing desk, and some LED lighting for ambiance you might have a hard time, right? After all you’re not an expert in building desks with all those additional features that have to work well together.

The expertise that people have in their chosen field cannot be understated. The time and effort that professionals put in when it comes to learning the regulations, analyzing the data, and making expert recommendations is beyond useful to those that need it. We go to doctors and lawyers for their expert opinions. Likewise, when it comes to IT solutions you need to consult with the experts as well.

The Antenna ExpertsVentev is a company that is experienced in the field of wireless antennas and power. They have spent years building solutions for organizations that need something more than a standard omnidirectional antenna for a wireless access point. Perhaps you’re trying to cover the seating at a stadium. You could be trying to deal with a room that has solid ceilings that can’t be modified. Perhaps you’re looking to deploy the new Wi-Fi 6E standard when it is ratified for use by the FCC and you don’t know the first place to start.

Ventev has spent thousands of hours building solutions for these problems. They know exactly how to engineer an antenna to operate in all three spectrum bands available to wireless professionals today. One antenna that covers all your needs. They also have specialized solutions to address things like the solid ceilings. Ventev can give you a way to conceal the equipment in a floor tile to make the wireless coverage flow up instead of down.

Image provided by VentevThe Challenges of IoTIn a recent interview with Jared van Allen, I talked to him about the solutions that Ventev has for one of the fastest growing segments of the market. IoT is more than just thermostats and smart toasters. Industrial IoT is one of the biggest segments of Venter’s customer base. Their solutions for a variety of hostile environments are unmatched. It starts with the understanding of how to integrate systems.

The Aruba CX4100iis a rugged IoT access switch with PoE+ and 60-watt PoE ports to run all manner of devices, from access points to cameras and more. It can operate in a variety of environmental extremes and offers you the capability to extend your network right to the edge to bring industrial IoT devices online. However, the switch is only part of the story. As Jared points out, if you buy a switch and an access point you have an angry customer because you don’t have a way to power those devices.

Ventev has an integrated enclosure solution that offers a cooled NEMA 3R box to contain your new Aruba switch. The enclosure also has AC and DC power as well as an integrated UPS to ensure that your devices stay powered during a potential outage. More importantly, the CX4100i and the UPS can communicate and use intelligence to know which ports of the switch to disable to preserve runtime for connected devices. You can tag certain ports to shut down in the event of a power loss first and preserve other ports with critical machinery as long as possible. This ensures flexibility and can keep your remote devices running longer than a less capable system.

The entire unit makes it easy to install as well. You can configure the switch on the ground by hooking up the box to power before climbing up to mount it. As someone that has found themselves typing in CLI commands from the top of a shaky ladder I can appreciate the simplicity of being able to do it from the ground to allow someone without a fear of heights to scale that ladder and get your device mounted and operational. Add in the IP68 connectors for your connections and you can see how the Ventev enclosure is a much more serene environment as compared to the typical hostile atmosphere found on a shop or manufacturing floor.

Image provided by VentevBringing It All TogetherVentev has the experts and the expertise to help you with your challenges. They know how to put together a winning combination of antenna designs, power, and more to solve your wireless and IoT needs. They don’t just start building without doing the research and providing the data you need to address your unique concerns. When it comes to projects that challenge your knowledge don’t be afraid to engage the power of those that know.

For more information about Ventev and their lineup of IoT solutions, including their integrated IoT solutions, make sure to check out their Aruba product page.


© Gestalt IT, LLC for Gestalt IT: The Power of Expertise at Ventev

View Details

In 2021, public records show that one in five businesses in Canada was hit with an information security breach. Larger businesses were most often targeted at 37% of the reported incidents, followed by medium businesses at 25%, and small businesses reporting at 16%.

If we break these numbers down, large businesses make up 0.2% of the reports at 2,936 organizations, medium businesses make up 1.9% at 22,725 and small businesses come in at 97.9% with a count of 1.2 million.

Of the $10B spent nationally in 2021 on information security remediation, almost half of that ($4.4B) came from large businesses, $2.9B from medium businesses, and the smallest amount, $2.4B from small businesses.

Put these together and we get a disproportional 62% of information security incidents targeting the top 2.1% of businesses. Budgets allocated to remediation follow this pattern with almost three-quarters of the money spent coming from that same sector.

Large organizations recognize that they’re a major target and invest appropriately. Smaller businesses however tend to look at these same statistics and conclude that they’re not significant targets, failing to take steps to protect their assets as a result. This turns them into even bigger targets in the process.

While it is true that the largest entities out there are more likely to experience an incident, the need for cybersecurity is universal.

BrickStor SPRackTop presents BrickStor SP (Security Platform), a cyberstorage solution that functions as a high-security NAS by front-ending block storage and providing SMB and NFS access through a high-speed access/analytics/audit/encryption engine.

Rolling SnapshotsA rolling snapshot of every file on the system is taken every minute, allowing for easy rollback to a known-good state.

When data is compromised through ransomware attacks and the like, it is relatively easy to recover the file and also examine what changes were made as part of a forensic investigation.

Professionals who are regularly involved in ransomware recovery know that undoing the rapid encryption process of a ransomware attack and detecting internal data modifications that are out of the normal patterns, such as hiding embezzlement, changing academic results, and are are two very different things.

Active DefenseThe core strength of the BrickStor SP platform is high-speed monitoring and remediation of unusual activity with the filesystem. Normal activity is constantly analyzed and factored into its learning engine. Abnormal activity is flagged and halted in short order, generating an alert to be dealt with by appropriate staff and/or processes.

Data exfiltration, where files are being copied rapidly from the storage system to local devices, is caught as soon as it goes over the system’s threshold for “normal” activity. High-speed data modifications, such as those found in ransomware attacks, are stopped in their tracks.

The Tech Field Day Showcase demonstration showed a ransomware process known to be able to encrypt files at speeds in the range of hundreds of thousands of files per second was stopped after changing only four, which could be recovered via the rolling snapshot functionality.

One Size Fits AllOne of the refreshing aspects of BrickStor SP is its market placement, or lack thereof. The product can easily scale from being a single VM front-ending on-board storage from a server’s directly-attached disks, to RackTop’s own SAS-attached hardware offerings, to front-ending large block-based SANs. Whether we’re coming into the game as a small organization or a large one, there’s a fit.

For the small organization with limited oversight capacity, BrickStor SP can be left to its own devices, providing data encryption, active defense and reporting. Larger organizations can customize this to meet their own processes such as requiring approval before active defense mitigation steps are taken. The administrative overhead of running the system can be as little or as much as is required by the business.

Regardless of how hands-off or hands-on the deployment, BrickStor SP learns from user patterns and provides an audit trail for post-incident forensics.

A Thought on “Cyberstorage” as a Term“Cyberstorage” initially seems like a strange word because the prefix “cyber” traditionally has nothing to do with security. It’s all about maximizing human intent by leveraging machine processes, for better or worse. Cybernetics, cyberbullying, and Kubernetes all fit into this. Modern definitions broaden this to include anything to do with machines or the Internet, but this still doesn’t have anything to do with security, so what’s with the term?

The major threat that the BrickStor SP solution is addressing is the one presented by human elements being influenced for improper access to information. Traditional information security is good at addressing who should be accessing something, but it tends to assume that the weak human link is much stronger than it is, or at least that there’s little to be easily done about it. The RackTop approach gets into the patterns of how they should be accessing things and takes action based on suspicious activity regardless of the permissions granted. It’s taking the zero trust concept beyond the individual and factoring in the activity itself.

Cyberstorage actively polices human/storage interaction and may well be one of the more aptly-named “cyber-” technologies out there.

The Whisper in the WiresMost organizations don’t have a team of people managing their cybersecurity policy and response. Many are lucky to have a team managing IT at all. The BrickStor SP product lends itself well to either, which is a refreshing turn in an industry that tends to focus on the top of the market.

Smaller companies that need a fire-and-forget solution that can prevent a breach in real-time and report the incident will find exactly what they’re looking for in this solution. Conversely, those who have large data sets with complex and granular access policies and controls will appreciate BrickStor SP’s flexibility to adapt to these. Best of all, there’s an easy migration path within the platform as the business grows.

Pricing is a discussion to be had with RackTop’s sales folks, but it’s well worth having that conversation.

Watch the RackTop BrickStor Showcase and more on the Tech Field Day website for more information.


© Gestalt IT, LLC for Gestalt IT: Defending Against the Widening Threat Landscape with Racktop BrickStor SP

View Details

As companies move to secure their digital estates from ransomware attacks leveraging zero trust architectures and defense-in-depth strategies, one asset looms larger than most – data. But in the fight against ransomware, focusing solely on data recovery is a losing game. The question one should be asking is whether to use an identity-centric solution, or a data-centric one?

With RackTop BrickStor, we can confidently say both.

Indentity-Centric or Data-Centric Security?The move to cloud-native applications, distributed teams, and unmanaged user devices has created a world of deperimeterization where IT estates span networks, devices, applications, data, and users across multiple platforms. These platforms are often built on infrastructures that companies neither own nor completely control.

If, perchance, they do control the infrastructure, they don’t always control the access that employees, partners, and customers have to resources over the internet.

The facts of deperimeterization have led to widespread acceptance of the need for zero trust architectures, such as described in NIST SP 800-207. Zero trust is a concept that moves us from providing implicit trust based on network boundaries (a secure perimeter) to a least-privilege model focused on explicit authentication and authorization of both subject and resource. This need for discrete authentication and authorization brings an obvious focus on identity.

Identity-centric security is a popular response to the demands of a zero trust architecture. It includes not only the identity of human users, but also that of data, devices, networks, workloads, and the interactions between all of them. Identity-centric security typically expands the traditional “triple-A” of authentication, authorization, and accounting, to include access policies and auditing capabilities.

But what is it actually trying to protect? In many cases, the most valuable assets being secured are various types of data. This has led to the emergence of data-centric security.

According to Wikipedia, data-centric security is “an approach to security that emphasizes the dependability of the data itself rather than the security of networks, servers, or applications.” Users/accounts also belong to that list. And it typically adds encryption and digital rights management to the conversation around zero trust architectures.

Why Not Choose Both?What if companies could have both identity and data-centric security? Enter RackTop Systems, and their active security platform.

RackTop Systems is a leader in the emerging “cyberstorage” market. Founded in 2010, they have a team of U.S. Intelligence Community veterans and engineers that are the brains behind their signature product, BrickStor SP.

RackTop describes BrickStor as an “active security platform for unstructured data” and claims it to be the only cyberstorage solution that addresses all five functional areas of the NIST Cybersecurity Framework, namely, identify, protect, detect, respond, and recover.

Cyberstorage is in many ways a direct response by the IT community to the seemingly overwhelming threat of ransomware and more recently, extortionware. Cyberstorage techniques also guard against the age-old “insider” threats of data exfiltration and destruction.

How does RackTop do cyberstorage? I recently joined a Tech Field Day Showcase to find out.

BrickStor CyberstorageBrickStor provides an excellent combination of data-centric and identity-centric security technologies to implement a zero trust architecture for defending unstructured data. It leverages encryption – FIPS AES-256 to be exact – and does not require any additional products to manage keys including policy-based key rotation and auditing.

A recently published technical validation paper from Enterprise Strategy Group says that “the RackTop software-defined data storage solution demonstrates robust proactive data security capabilities in an all-in-one Cyberstorage package that is easy to use.”

Even more exciting is the combination of UEBA (user and entity behavior analysis) and SOAR (security orchestration automation and response) capabilities that are purpose-built around the realities of securing data.

With its well-tuned, out-of-the-box policies, BrickStor can start protecting the most valuable data right from day 1. To see this in action, check out this live demo.

The Bottom LineIn order to secure organizations against the complex and evolving threats, a zero trust architecture that is both identity-centric and data-centric is critical. Luckily, RackTop Systems has done all the heavy-lifting by embedding these techniques and technologies into their BrickStor cyberstorage platform, bringing to customers a ready-to-use solution.

Watch the video from the RackTop Showcase on the Tech Field Day website for more information.


© Gestalt IT, LLC for Gestalt IT: Identity-Centric and Data-Centric Security with RackTop BrickStor

View Details

As connectivity of environments has evolved over the years, cyber threats have proliferated making proactive security fundamental to good cyber health. The secret to averting security crisis lies in organizations’ ability to anticipate potential situations and adopt preventive measures ahead of time.

Recently, I took part in a Tech Field Day Showcase with RackTop Systems where proactive security measures was a key theme. RackTop Systems is changing cyber protection with BrickStor SP, a cyberstorage solution that has the ability to defend data assets against ransomware and other malicious threats.

Zero Trust at the FoundationRackTop Systems takes a transformative approach to storage security. It leverages a zero-trust architecture that is built on the core principles of authenticate, authorize, and validate. The model works by eliminating implicit trust and replacing it with continuous validation before granting access to entities.

Designed with a data-centric zero-trust principle, BrickStor observes every interaction that’s happening in the environment, whether or not it involves suspicious users. So even if an authorized user attempts to make an unauthorized move, it flags the transaction as a security alert.

BrickStor SP provides active defense against security breach, data exfiltration and insider threats through real-time detection and diagnosis and quick response and recovery. The solution is light and agentless, and can be installed under 15 minutes. Ransomware mitigation starts almost immediately after it is deployed.

Key CapabilitiesBrickStor is a software data storage solution that stores, manages and protects data. Designed with advanced security and compliance features, it is a one of a kind storage solution that focuses on the most burning cyber issues like ransomware attacks and insider theft.

Key capabilities include real-time monitoring and detection, alerting against unusual and suspicious activities, and user and entity behavior analysis (UEBA). RackTop uses high-level encryption to protect data at rest and in flight. Inside the storage, data is encrypted at both physical and logical layers. Built-in policy-based key rotation and auditing eliminate the need to use secrets manager.

BrickStor SP has an intelligent UI that displays details of the breach, including where the issue began and at what stage it was contained and quarantined. RackTop’s demo highlighted these areas and illustrated policy-based data protection. This includes WAN-optimized replication at the cloud level, or with a physical NAS solution. Among the key features demonstrated were compliance-based reporting and active defense that enables tracking of suspicious behavior at different levels. Administrators can assign policies and watch who’s accessing data and at what time.

Automation & OrchestrationOne of the key features of BrickStor SP is its ability to leverage automation. As the threat landscape changes, not every organization will have the resource or bandwidth to fully dedicate a security and incident response team. RackTop’s solution with its simplistic deployment and automation capabilities is a life-saver for them.

The automation piece allows it to be scaled to any level within the organization. Storage performance is unaffected by the dual encryption. IT professionals can define policies and automate, letting the solution work on auto-pilot.

Real World UsesRackTop has been instrumental in mitigating ransomware and other cyber threats since the early years. With data provided in real-time, post-attack forensics enable users to dive into the processes and events correlating to the attack, making incident response at file-level quick and effortless.

Many different industries from state and local governments, public sector, to legal entities have leveraged BrickStor SP. It is deployed in healthcare facilities where it automates data protection. Flexible retention policies ensure that data is protected at all levels, and backed up and replicated in compliance with HIPPA and PCI measures.

ConclusionIt’s safe to say that RackTop sets a new standard of data protection with BrickStor SP. Its zero trust architecture ensures that data remains confidential at source, while capabilities like auditing and real-time monitoring of file access and high-level encryption guarantee the highest level of protection against data breach.

Watch the full RackTop Showcase on the Tech Field Day website or YouTube Channel. Read more about RackTop’s CyberStorage for Active Defense on their website.


© Gestalt IT, LLC for Gestalt IT: Active Defense Against Ransomware with RackTop Systems’ BrickStor SP

View Details

Modern storage is a complex combination of both hardware and software. While software-only storage options do exist, if you’re after high-end performance, you’ll have to deal with the nuances of hardware sooner rather than later.

And this is especially true when dealing with modern flash storage. Assumptions from the HDD era no longer apply. What works well for spinning rust is often irrelevant, if not outright harmful, when applied to modern flash.

Software Runs on HardwareTalking directly to flash as it wants to be spoken to removes friction that only exists because of outdated abstractions. A SATA bridge between the outside world and the flash devices was useful when systems needed to speak SATA to the storage. But what if we don’t?

If we move to a higher level abstraction such as iSCSI, SMB or object protocols, why keep using an outdated bridge to nowhere? Translating from S3 to SATA only to re-translate to flash just slows things down. Why not skip the SATA bridge completely? That’s just what Pure Storage did with its DirectFlash approach. Instead of using commodity SSDs, Pure uses flash media directly.

To do this requires knowledge of what the flash hardware layer is doing, and what it is capable of. Changing the software in ignorance of the layer below the abstraction will only lead to pain. But with expert knowledge of how things really work at the hardware level, software can be made to dance. Consider the ways in which the early Apple II was designed to provide color at a time when color monitors were rare and expensive. And yet it required fewer chips to achieve!

These kinds of advances require working with both hardware and software together, rather than in isolation. Software runs on hardware, and understanding the target environment well can provide an edge to those looking for superior features or performance.

Yet as consumers of storage, we would usually prefer not to care too much about how things happen, so long as they do. We want to use abstractions, not talk directly to the flash ourselves.

The Right Amount of AbstractionThe challenge with abstraction is finding the right amount to use. Too little, and we have to wrestle with too many trivial or frustrating details from the underlying systems that were supposed to be abstracted away. Too much, and we’re unable to get our work done. Neither is appealing as a customer, and these issues often reveal themselves at the worst possible time. They’re exposed by edge cases, by things that aren’t part of the day-to-day.

Storage vendors like Pure encounter these edge cases early, often as they’re developing new product features and seeking feedback from customers. They get a variety of perspectives in order to figure out what the right amount of abstraction is in practice, not just in theory.

Yet the ‘right’ amount of abstraction is hard to prove objectively. Most of the time it’s a matter of judgement, taste, and wisdom earned from experience. While there may be some customers for whom manually tuning the compression algorithm or data layout might make sense, most are better served by letting Pure automate things.

By knowing the internals of how and when data is written to the flash hardware, Pure can adjust how the software functions on the fly. Data can be moved, or not moved, depending on the condition of the flash. Compression algorithms can be adjusted to suit the workload, automatically.

Freedom from FrictionHaving both hardware and software options available opens up more ways to address design challenges like this. Is a new compression algorithm best implemented in hardware, or in software? Should it be custom hardware, or a commodity part?

Or perhaps, as with the recently announced DirectCompress accelerator card for FlashArray//XL, a combination of software and hardware works best? DirectCompress uses a fairly standard component—an FPGA card—and combines it with Pure’s software expertise to improve inline compression performance. Other choices were available, like ASICs, GPUs, or staying with CPUs, but Pure’s experience and judgement guided the decision.

As customers, we can happily stay using abstractions and concentrate on the results. Does using DirectFlash improve performance? Does an FPGA help our workloads? We have performance and functionality goals to meet, so what ultimately matters is if our storage choices work.

Pure’s expertise with both software and hardware working together means it has more options to succeed, and so do we. Read more on FlashArray//XL on the Pure Storage website.


© Gestalt IT, LLC for Gestalt IT: Innovating with Both Hardware and Software with Pure Storage

View Details

I started creating server build automation during a change freeze, immediately after New Year’s 2000, the Y2K bug apocalypse date that didn’t quite eventuate. Ever since that start, I have much preferred automated build processes to relying on people following a list of manual steps. One of the reasons for that is the principles of IT infrastructure automation fit nicely with the DevOps culture that enables faster application development.

I create build automation designed to have a human watching over, mainly watching for errors. Building automated error handling and verification is complicated and time-consuming. Watching that build process is acceptable when building a small lab. It was also great when three or four engineers built a couple of dozen servers for doctors’ surgeries.

The problem is that human operations are costly to scale, particularly to many locations spread over a wide geographic area. This is what is defined as the far-edge location. These locations might be retail stores, or delivery trucks, often places where business staff do their day-to-day work, interacting with the customers and directly generating revenue.

Increasingly, technology is either making the staff at these sites more productive, or adding ways to generate more revenue from the customers that visit these sites. But this technology has a cost, and business owners need to control those costs to turn that revenue into profit.

Control Far-Edge CostsOne of the important ways to control IT costs for the far edge is to ensure that one never needs to send a skilled IT engineer to sites. Everything that goes to the sites should be able to be installed and connected by the onsite staff, whether they are the driver of the truck or a checkout supervisor at the store.

Once the hardware is physically installed, every other part of deployment and management must happen from a central console where the IT team can manage groups of sites and remotely resolve faults. Ideally, the platform used should automatically fix as many faults as possible, which brings us neatly to Scale Computing and its background in self-healing computing infrastructure.

Zero-Touch ProvisioningScale Computing has always designed SC//HyperCore HCI to look after itself as much as possible. SC//Platform extends to edge deployment and is the basis for Zero-Touch Provisioning (ZTP) of SC//HyperCore clusters at a grand scale.

ZTP was the centrepiece of Scale Computing’s recent presentation at the Edge Field Day event in San Francisco. In the live-streamed presentation, four clusters were deployed from a central console to five Intel NUCs. The “onsite” deployment was connecting power and ethernet to the appliances. Everything after that was achieved using the SC//Fleet Manager console, which had been pre-populated with the hardware IDs of the physical nodes.

In actual deployment, these hardware IDs are harvested from the fulfillment system that delivers the appliances to the site. SC//Fleet Manager is easily pre-populated with the IDs before the devices are delivered to the site. Once the appliance is powered up, it phones home to SC//Fleet Manager over the Internet and appears in the customer’s view of the console.

Application PlatformsSC//Fleet Manager handled a lot of configuration and initial setup for users, deploying the latest version of SC//HyperCore to the nodes, and deploying a set of standard initial VMs. The initial VMs might be the standard edge site servers with locally installed applications, or home to cloud-managed applications using platforms such as Google Anthos or Azure ARC-enabled applications.

SC//Fleet Manager can drive updates of SC//HyperCore to the clusters, and Scale Computing has an Ansible collection for managing multiple clusters over time. The Ansible collection is Red Hat certified and available on Galaxy for easy installation.

Scale Computing Removes ProblemsScale Computing’s approach of removing the burden of routine tasks and fault remediation from customers is widely liked. ZTP is an excellent part of continuing the story where Scale Computing can deliver large numbers of cost-effective clusters, or single nodes, to deployment at the far edge.

The Intel NUCs, using specific models that incorporate feedback from Scale Computing, make a great compute node to deploy in far-edge locations. Scale Computing also has its data centre-specific models for less constrained environments.

Scale Computing showed its commitment to partners by having Avassa and Mako Networks present at the Edge Field Day event in their offices.

Be sure to watch the Scale Computing presentations from the recent Edge Field Day event, and keep up with what the other delegates think at TechFieldDay.com.


© Gestalt IT, LLC for Gestalt IT: Scale Computing Removes Problems at the Edge

View Details

Last week, I met with the Pure Storage’s FlashBlade product team to hear about their newest offering for unstructured data storage. Given Pure’s focus on high-performance, scale-out all-flash solutions, I assumed that I’d be hearing about a new edition of FlashBlade//S. But instead, we talked about their new software-plus-hardware product aimed at workloads with similar access performance needs: FlashBlade//E.

The Current Storage MarketTypically, oragnizations keep their cooler data sets stored on spinning disks or other non-flash options to save on costs. Although they tend to be slower, these technologies remain the cheaper alternative for large data sets that are very infrequently accessed.

Traditionally, data professionals think of data access needs as hot, warm, cold, and archive. The acquisition and operational costs decline from hot to archive.

However, until now, the most cost-effective option for cold and archive data is spinning disks that are available at affordable price points. There is hardly any organization left that uses tape backups, for these reasons.

Of course, nothing stops an organization from hosting all its data on high-performance infrastructures other than the fact that data growth continues to be a significant factor contributing to the total cost of ownership for storing more and more data. Additionally, most of these cooler datasets tend to grow faster than transactional data. That runs up the cost really really quickly.

FlashBlade//EPure’s new solution builds on the same software and hardware architecture as the FlashBlade//S for unstructured data. Pure is delivering this offering with nearly all the features of their premier product while cutting costs to the point where organizations can save money moving from non-flash options.

Same FeaturesPure Storage wishes to make FlashBlade even more accessible and convenient for their current and new customers. To that end, it has included everything that FlashBlade//S solutions have:

  • Scalability: Hardware uses the same extensible and upgradable Direct Flash Modules
  • Capacity: Up to tens of petabytes
  • Networking: Up to 16x100GbE, doubling max of FlashBlade//S
  • Energy: 2,300W (Compute plus Storage) or 1,600W (Storage only)
  • Software: Runs the same Purity/FB software
  • Size: Hardware dimensions are much smaller than disk storage
  • Data Security: Supports the same features such as SafeMode immutable snapshots as well as always-on encryption

The datasheet for FlashBlade//E has more technical specifications.

Lower Cost Lowering costs is the primary focus of this product. Pure says that the acquisition costs are similar to those of comparable disk solutions, and at a total cost of ownership of about $0.20 US per GB raw, it is 40% lower. Pure also points to costs savings over traditional storage because all-flash saves on energy use, at about one-fifth of that of disk systems.

Target WorkloadsPure uses the term repository workloads to describe the best fit for data access and usage for FlashBlade//E, and I think that helps explain the types of workloads they want to support. These include traditional unstructured data such as media, backups, snapshots, data lakes, documents, and records. These are large volume datasets that have a huge variety of performance and access needs.

Sample Use CasesConsider medical imaging as a use case. It would involve storing digital artifacts like these at a high resolution for a very long time. But in reality, they get accessed only a few times a year, depending on how long ago they were created.

The same goes for large data lakes, where data always flows in but never leaves the lake. Moreover, given that data lakes typically support data in its raw, non-aggregated form, they tend to be loaded with mostly uncurated data, and have no set structure to the same data from multiple sources. Add to that the fact that we often store raw, cleansed, and curated versions of these data. Data lakes grow faster than transaction systems such as data warehouses.

IoT streaming media is another great example of large volumes of data being stored in raw, cleansed, and analyzed versions. You might need to process all that data at the time of ingestion and archive it to storage for the longer term.

Target CustomersPure wants to entice their existing customer base to migrate all their spinning disk storage to FlashBlade//E, and that makes perfect sense. Expanding footprint in a datacenter is always easier if you already have products there. So if Pure’s estimates for cost and energy savings come anywhere near what they are predicting, it’s a win-win for both organizations.

However, Pure also sees a path for new customers with different requirements for high-performance all-flash storage. With the new price points and cost savings, these organizations with large unstructured data needs will see sense in considering FlashBlade//E as a solution for their repository workloads.

What Does the //E in the Name Stand for?As you may have guessed, the main focus of the E branding is Environmental. However, Pure also points to the lower cost feature as Economical, the simplicity provided as Effortless, and their Evergreen//Oneoptions as Everlasting.

With 10 to 20 times greater reliability and innovative designs, Pure claims that their products deliver 85% e-waste reductions over traditional disk storage. Given all this, I think this product deserves the //E branding.

Karen’s ObservationsPure has created a new entry point for organizations that have not yet transformed to an all-flash storage. With FlashBlade//E, they have opened a door for current customers to move to a more sustainable, energy-efficient set of storage solutions. I’ve previously written about Pure Storage’s commitment to building products that reduce energy usage and mitigate planned obsolescence. Their scale-out framework enables companies to extend the life of their storage infrastructure easier and with less downtime. To quote from my previous article, “This is the way I want all technology to go: upgrade instead of replace.”


© Gestalt IT, LLC for Gestalt IT: Meet FlashBlade//E, Pure Storage’s All-Flash Solution with Same Features but Lower Cost

View Details

As organizations embrace edge computing, they seek the ability to deploy technology to deliver new capabilities to business operations, and open new avenues for revenue generation. So far, the focus has been on the delivery of infrastructure for static workload management. But a long-time goal of edge adoption has been to dynamically control workloads from afar, making application updates, deploying new services, and keeping edge points secure both feasible and cost and time-efficient.

Stepping UpScale Computing, a long-time leader in HCI infrastructure for a variety of edge environments, has had great success in deploying edge solutions.

Amassing thousands of customers during their fifteen-year history, their reputation for delivering solutions that work out-of-the-box has earned them over 100 points on CRN’s vendor product quality score, a historic high for the publication.

As the edge has matured and customer requirements have become more advanced, Scale Computing has upped its game with software development to ensure a full lifecycle of remote oversight of infrastructures. It’s in this work that the company rests its opportunity for continued market advancement as the momentum of edge proliferation grows.

Full Lifecycle ManagementThe foundation for this full life cycle management is Scale Computing’s Autonomous Infrastructure Management Engine (AIME), custom-built software that sets Scale Computing deployments apart.

AIME is built deeply into the hardware telemetry to oversee everything from BIOS and firmware health to hardware states, environmental measurements, network state and more. With always-on oversight, Scale Computing is aware of infrastructure issues well before its clients, and can timely trigger remediation.

When you consider the diverse landscape of edge environments and the relative skillset of edge-based employees, this capability is a gift to organizations, to deploy with confidence, and mitigate risks of truck rolls for hardware troubleshooting. AIME’s sophistication also ensures that data is properly mirrored, and any impact of loss is minimized.

Scale Computing has built on hardware control achieved in AIME with the newly released zero-touch provisioning (ZTP) in SC//Fleet Manager. ZTP allows Scale Computing to pre-load software on edge devices so that setup at the edge is simplified to plugging equipment to the location. This is when SC//Fleet Manager takes over and enables remote initialization of each cluster, deployment of initial configuration, and application of the cluster name.

The combination of SC//Fleet Manager and AIME also de-risks initialization errors such as a device plugged into a wrong location with seamless re-boot and re-provisioning of the cluster. For organizations wishing to deploy fleets of edge devices in environments with limited or no local technical support, the cost savings that ZTP offers expand the opportunity for edge device deployment.

With ZTP and AIME, Scale Computing has achieved initial deployment and cluster oversight of its edge solutions. But customers want more. They want the opportunity to update software across their edge fleet as efficiently as possible, and for this capability, Scale Computing has developed their SC//HyperCore Ansible Collection.

These Red Hat certified open source-based playbooks are provided to enable IT administrators to provision clusters, deploy applications, and update security and compliance with central, automated control.

This technology’s elegant approach extends the value of edge infrastructure by enabling organizations to deploy new services on existing hardware, and extending the value of edge deployments for business evolution. It’s also reflective of the broader direction of edge computing delivering more sophistication of workload opportunities to monetize data at the edge.

ConclusionOrganizations that seek an edge infrastructure provider with seasoned experience deploying fleets of equipment across edge environments should seriously consider Scale Computing’s offerings. Its software portfolio reflects a deep understanding of the unique capabilities IT organizations require for remote site deployment. Its engineering investment in hardware telemetry and software oversight ensures that central IT operations can effectively oversee remote deployments. Lastly, its investment in SC//HyperCore Ansible playbooks makes sure that organizations are not limited to appliance like edge devices. Scale Computing also lives up to its name in being able to support a broad scale of devices and types of configurations to suit a wide array of workload and environmental requirements.

To learn more about the company’s offerings, visit Scalecomputing.com and request a free demo, or check out the presentations from the recent Edge Field Day event at TechFieldDay.com.


© Gestalt IT, LLC for Gestalt IT: Achieving Dynamic Provisioning and Control at the Edge with Scale Computing

View Details

Often, engineers are sold a shiny new technology that’s positioned as the answer to “everything.” And just like that, the technology that had been serving us well for years is suddenly stigmatized as legacy technology and there’s the urge to rip it out and replace it. The reality though, is that we generally don’t have the luxury of rearchitecting an organization’s entire infrastructure, nor should we.

The better approach is looking into how these new technology innovations can help make the current architecture better and address specific use cases more effectively. SD-WAN and the emergence of SASE is one such example of technology innovation that allows organizations to apply secure access no matter where their users, workloads, devices, or applications are located.

Single-Vendor SASEWhile some probably remember when SASE stood for “Self-Addressed Stamped Envelope,” and was used to get something from a celebrity or some back-of-the-box offer, today it stands for Secure Access Service Edge. Now wait, this is old stuff from 2019, right? While the term has been around for a while, the technology really matured as the remote work population reached all-time highs throughout the pandemic.

With the majority of organizations supporting hybrid workforces, CIOs are tasked with the challenge of securing users as they move from their home to the office, during travel, and everywhere in-between. If it wasn’t dead before, the perimeter for securing a business is officially gone now. Zero trust is the only way forward, and SASE is a major part of that future state. SASE is not a silver bullet, no technology is, but it plays a very important role in securing a distributed workforce and should be used to augment your next-generation firewalls and other on-premises technologies to help provide consistent security and user experiences regardless of where employees are located.

This is where Fortinet’s approach with Single-Vendor SASE comes in. Customers want FortiOS to be the unifying presence, from firewall to the cloud and beyond. Whether it is on-premises or in the cloud, FortiOS is the glue. Likewise, one client, FortiClient, will help secure endpoints whether they are connected to the cloud, or a traditional on-premises firewall.

Unified policies, whether it’s “legacy” technology, or the latest iteration of SASE, provides business with a way to seamlessly integrate the new technology while maximizing their existing IT investment.

Another reason for extending existing security architectures instead of replacing them is that, not all use cases lend themselves to SASE. For example, many organizations still demand on-prem security for use cases like internal segmentation, compliance and regulatory requirements – especially those in the financial and federal/government verticals. And, of course, there are customers are cloud-friendly or predominantly remote workforce would take the route of cloud-delivered solution. What is turning out to be more important is the flexibility for customers to enable security where it is needed on their journey to a SASE framework and, while these organizations are adopting cloud-delivered security for a hybrid workforce, they are also making sure to enable internal segmentation for on-prem locations to prevent the lateral movement of threats.

In this scenario, the single-client approach from Fortinet shines as it reduces the individual policies, configurations, and software that need to be deployed to the end user devices, thus simplifying deployment, as well as user experience. They use the same software for both on-premises and SASE. So, to them, it’s the same thing and it’s simple. No need to have complicated instructions of when to use one solution over the other. Colleagues outside of IT deserve connectivity and security with the least amount of friction possible.

ConclusionWhile it may not always be possible to use the same vendors for all security needs, there is a compelling case to consider it. Keeping a uniform security posture regardless of where users are helps make sure that security intent is properly implemented. Misconfigurations and oversights are often at the root of security infiltrations. Anything that helps IT teams provide a great user experience while maintaining a uniform security posture is a win-win for the business.

To learn more about what Fortinet is doing with their Single-Vendor SASE architecture, be sure to check out their page.

GuestNarav Shah, Vice President of Products – SD-WAN, SASE, Zero Trust at Fortinet

ModeratorBen Story

LinkedInConnect with Narav on LinkedIn and learn more about Fortinet and their security on their website.

Twitter@Ntwrk80

Follow us on Twitter and SUBSCRIBE to our newsletter for more great coverage right in your inbox.


© Gestalt IT, LLC for Gestalt IT: Fortinet – Securing a Hybrid Workforce Does Not Mean Rearchitect Everything

View Details

Edge computing continues to grow as demands for low-latency, disconnect-friendly, privacy-conscious applications swell. But even though edge computing accommodates new and creative architectures, it does not come without its costs.

Edge environments face numerous constraints that are foreign to the cloud computing world: brittle network connections, poor bandwidth, unreliable power, constrained physical spaces, and limited hardware resources to name a few.

This problem domain is not foreign to me as we navigated similar architectural challenges while provisioning and clustering 2,500+ K3s clusters on bare metal Intel NUCs across all Chick-fil-A restaurants in North America.

The Challenges of ProvisioningOne of the many challenges with edge computing that is easily forgotten is the absence of on-site technical support, especially in remote locations (think oil fields), or retail environments.

Enter zero-touch provisioning (ZTP), which enables devices to be automatically provisioned with the necessary configuration for their environment without any human interaction.

CNCF defined this key principle for edge-native solutions in a recently published Edge Native Report as: “Edge native requires a mix of remote and centralized management and zero touch provisioning of hardware and software. Staffing at the edge may be untrained, untrusted, minimal or even non-existent”.

Organizations considering large-scale edge deployments will likely find ZTP essential to their success, since manually provisioning hundreds, thousands, or tens-of-thousands of devices is too slow, cost prohibitive, and mistake-ridden to do with a human team.

Achieving zero touch provisioning requires solving several challenges. Devices must be imaged and assigned to customers, shipped to the correct locations, and tracked throughout provisioning. Device trust must be established.

Provisioning failure states must be considered, accounted for and managed, including scenarios such as loss of power or the network during provisioning. Most importantly, “bricking” (the process of turning a functional computer into a non-functional brick) devices—and especially fleets of devices—must be avoided at all costs.

Zero Touch Provisioning with Scale Computing At the Edge Field Day event, Scale Computing announced the launch of their Zero Touch Provisioning capability which provides a solution to this. A customer can select a device from the available fleet, that includes everything from lower-capacity devices like Intel i7 NUCs with 64GB RAM, to powerful machines with GPUs, providing a lot of flexibility on compute power and physical form factor.

Scale Computing images and ships these devices to the site, where they simply need to be plugged into power and ethernet. From there, Fleet Manager and the zero-touch provisioning take over.

Scale Computing’s architecture revolves around a proprietary state machine that is built into their HyperCore product, which is present on the initial image shipped to the site. This state machine tracks a node’s progress through its initialization process. This approach handles failure states well, and prevents in-field “brickability”.

Here is how it works. When a device is connected, it checks in with SC//Fleet Manager, which delivers a configuration file that tells the device what to do. Devices are pre-assigned to a particular customer before shipment to establish a basis for trust, and enable the devices to be discovered for configuration within the Fleet Manager portal.

Scale Computing doesn’t know what the customer wishes to do with the node at this time; they simply provision it and ensure that it initializes successfully. Provisioning is quick, taking around 10 minutes.

Once a device is initially provisioned and clustered, Scale Computing provides an Ansible-based capability to configure whatever is desired on the cluster of nodes, such as Google Anthos, Azure Arc, Avassa, or Kubernetes.

Cattle, Not PetsIn cloud computing, DevOps teams are encouraged to treat their compute resources as “cattle, not pets”. “Don’t bother with naming your servers or getting to know them, because they could be gone any moment,” they say.

While the edge development paradigm is quite different than the cloud, this principle still works. This is especially true for deployments with low-cost hardware, such as the Intel NUC (which was exceedingly popular at the Edge Field Day, and is available from Scale Computing).

With the right architecture in place, it is possible to treat edge nodes as cattle instead of pets, too. As Craig Theriac of Scale Computing stated, we could think of these edge devices as “disposable units of compute”. Should a problem arise, we simply “shoot the cow” and drop-ship a replacement knowing that the workloads/service remain available on the remaining nodes even while waiting for the replacement to be added.

Assuming ZTP is a capability, this makes a lot of sense in many edge deployments as shipping a replacement device is inexpensive and reduces operational troubleshooting toil and the need to deploy a human technician. Failed nodes can be shipped back, validated, and either, be re-entered into the fleet or disposed of.

Having a mature process for “wiping” a deployed node back to its original arrival state and re-creating it is another pragmatic way to avoid costly human troubleshooting. Manually supporting issues becomes problematic at the edge scale, and requires a cattle mindset, but it is only supportable with automation.

ConclusionEdge is a frontier, and at times a wild one. Remote locations, low-trust environments, unreliable networks, and data privacy laws are real challenges, but businesses and consumers demand these frontiers be conquered to create the high-quality experiences they desire. Zero Touch Provisioning and a “Disposable Units of Compute” philosophy are two approaches that can power the acceleration of successful, scalable and manageable edge deployments. Scale Computing’s Zero Touch Provisioning and Sc//HyperCore solution appear to be great building blocks that organizations can place at the foundation of this new chapter.

Be sure to check out Scale Computing’s presentations from the recent Edge Field Day event to know more.


© Gestalt IT, LLC for Gestalt IT: Scaling the Edge with Zero Touch Provisioning and Scale Computing

View Details

CXL is poised to revolutionize the datacenter. It is real and it is happening now! Micron is in the unique position to deliver CXL solutions in multiple categories to obtain the desired outcomes that CXL enables. The resources and wide memory and storage portfolio Micron has to offer reflects their industry experience and vision; together it creates a synergy for the CXL solutions being brought to market. If you will explore the problems CXL attempts to solve, the likelihood of the solution being a Micron-enabled option will seem logical.

What is CXL?Compute Express Link (CXL) is an industry-supported cache-coherent interconnect. CXL allows for the enablement of a disaggregated and composable datacenter. This modern interconnect will extend the server resources beyond the confines of a single form factor. CXL is based on PCI physical and electrical interface, but unlike past interconnects, it will address the composability and performance characteristics required for the category of compute resources being extended.

Cache-coherent, multi-point sharing and expansion are addressed with CXL. CXL is different from legacy expansion because at its core is cache-coherency for all the expanded resources. Limitations of a single form factor are no longer a concern, instead high-speed interconnections are enabling resources to be made available sans the limitations of form factor like power, heat, space.

The enablement of CXL will drive innovative compute design adoptions, not at the motherboard level, but it will open the motherboard to include rack design elements. An adjacent rack with all the available resources of more power and more heating and cooling options will be part of the new composable server footprint. Compute with CXL will be more than agile; it will enable pliability and you will not be hindered by a system CPU designed 10 years ago. CXL provides a clear and open standard that the entire industry is supporting to unlock innovation.

How can CXL Help? CXL will shift memory as it is now, in terms on how it is consumed. It will enable the next revolution in datacenter computing and Micron is poised to be a part of this. Micron is a US-based memory developer, manufacturer, and supplier with fabrication plants in the US.

Micron has announced expansion plans worth $15 Billion in Boise, ID and $100 Billion in Clay, NY. Micron is not sitting still. But what problem are they trying to solve? Compute is becoming denser as a two-socket system. It is not defined by the number of sockets but the number of cores available to each socket. That compute density inherently poses the problem of how much denser RAM must be and how much faster is needs to be. DDR5 attempts to address capacity and bandwidth, with more memory channels. The current trends show CPU development is outpacing DRAM, but the cost of stacking DRAM drives the TCO cost to be a very expensive option. Enter CXL.

CXL + Micron = What A Combo! The combination of CXL and Micron technologies are unlocking the much-needed expansion without the inherent penalty of performance loss. Current options of memory stacking make the linear cost-per-bit impossible. Current system memory (RAM) is 40 to 50% of the material cost of a server and the need for memory intensive workloads is not going away. Use cases like AI/ML demand more performance and that includes RAM.

CXL allows for memory to be a pooled resource on the PCI bus via CXL. A lower TCO versus costly TSV (Through Silicon Vias)— traditional wire bond that will ensure 40 to 100 Mega transfers per second. CXL while a little bit slower than DDR5 is still cache coherent. Cache coherence defines the behavior of reads and writes to a single address location and this is very important for the integrity of computation across multiple CPU and shared resources. Leveraging the PCI bus, CXL provides for resources to be available further away from the CPU and memory bus without performance penalty.

Other areas of computing alternatives can also be addressed head-on with CXL. For example, the system motherboard and channel bus limitations. The traditional Two DIMMS Per Channel (TDPC) ratio is channeled by incorporating more memory than was possible before. System complexity and balanced memory systems vs CPU is addressed. GPUs can be leveraged more effectively across multiple systems. These current challenges are addressed with the advent of CXL as a rapidly maturing option.

ConclusionMicron’s fabrication of memory (DRAM) and other CXL attached memory products will contribute to the next generation of computing. At the OCP Summit, Micron presented its position in the CXL revolution and value proposition to solve the challenges of modern datacenter computing.

At the OCP Summit Industry Interest in CXL was very strong. Hyperscalers and more.

The call to action is in front of us today: Be Prepared. CXL is alive, real and shipping. Future designs will leverage the benefits provided by CXL. The echoing message of TCO being lower will have to be vetted, as alternatives are constantly being challenged by supply change and physical limitations of current technology options. I look forward to innovative and future designs that incorporate CXL, and to seeing how existing technologies will incorporate CXL. Consider HCI and SSD supercharged with CXL. What will a single HCI compute be composed of? How will it easily be recomposed with CXL? What hasn’t been discussed, is how the hypervisor giant, VMware, will leverage the CXL advantage? The race of software development to catchup to hardware development and vice versa continues. Either way Micron will be there to drive the desired outcomes. Will you be ready?

If you want to learn more about what Micron is doing, Check out their Tech Field Day presentations. Micron Technology’s Ryan Baxter also guested on our Utilizing CXL Podcast, part of our Utilizing Tech series. Watch his episode to hear more on Micron’s CXL efforts. If you want to learn more about what Micron is doing, you can also check out their Tech Field Day presentations.


© Gestalt IT, LLC for Gestalt IT: Micron and CXL: OCP Summit Update

View Details

Introduction On October 20th, 2022, Micron took part in the Open Computing Project’s CXL Forum. The forum was a marquee part of the OCP’s Global Summit that entailed an entire day of research, describing plans for CXL, and showing off both current and future CXL technologies. Micron’s session, which was nothing short of excellent, discussed CXL’s market outlook in terms of just a few of the technology’s use cases.

CXL’s Capabilities, Past and PresentCompute Express Link (or “CXL”) is a developing standard for linking together devices via PCIe that previously only communicated inside of a server. The kinds of devices that can be linked depend on the version of CXL under discussion. For example, in 2019, CXL 1.1 introduced the ability to communicate directly from the CPU in one rack-mounted server to a memory expander located either in the server, or on another rack-mounted node entirely. CXL 2.0, introduced in 2020, added support for single-level CXL switch so that systems could connect to multiple CXL-compatible devices via a fabric instead of requiring 1:1 connectivity. For 2022, the CXL 3.0 release describes connections that can use PCIe 6.0, multi-level switching, and a double total bandwidth, among many other features.

CXL’s end goal is to completely redefine the data center to allow for greater flexibility in design. The standard is backed by every major manufacturer you can think of: from hardware and CPU manufacturers such as Intel, Micron, and NVIDIA, to software and platform providers such as IBM, Google, and Meta.

In 2021 it was announced that Micron had stopped developing 3D XPoint memory in favor of focusing on the CXL future. This looks to have been the right decision, as the presentation that Micron gave showed that the potential of CXL will greatly surpass what had been hoped for from the 3D XPoint project.

Micron’s presentation highlighted the fact that CPU designs have been placing an emphasis on higher core counts, and that chip and motherboard designs mean that standard memory deployments have not been able to keep up from a bandwidth perspective- especially for resource-intensive applications. What applications need is faster access to more memory, and CXL can provide that. This is accomplished because the PCIe bus handles the channels between the processor and memory. CXL is not limited by the number of DIMMs per channel limits of the motherboard – limits that will be more restrictive in the future, owing to new DDR designs.

Using CXL to Solve Memory Management Problems TodayCXL Memory Pooling is a feature that is designed specifically for this problem. If you have a host that needs more memory, you can assign that memory from another host in a dedicated fashion. This CXL bus connectivity will allow ultra-high memory allocations per host that far exceed the main system memory limits.

See the diagram below that shows how CXL 2.0 enables Memory Pooling:

It is true that the memory on CXL will be slower than DIMMs on motherboard, but not as much as you would think (Micron states that CXL-based servers show latency from CPU to CXL memory that is “equivalent to a single NUMA hop”). Micron also highlighted that the workload capacity and bandwidth can be ‘dialed in’ as system performance is observed and fine-tuned over time to take advantage of the different memory ‘tiers.’ There are a number of very memory-hungry applications out there, such as AI/ML, NLP, and in memory databases that will significantly benefit from this additional system memory capacity and bandwidth even if it isn’t 100%-line speed compared to on-board DIMMs.

ConclusionThere is a lot to like about CXL, and Micron showed the performance advantages with both the existing hardware and CXL’s future. Additionally, there is a TCO benefit to only upgrading the memory allocation via a CXL-compatible device compared to buying a whole new server. In the post-CXL 3.0 future, this kind of flexibility will extend from just memory to basically any system component you can think of – CPUs, GPUs, storage, and more.

While the standard has been around for a while (remember, CXL 1.0 was announced in 2019), it is only in the last year or so that we have seen real life examples of the technology in action. Micron’s presentation showed the practical benefits CXL will bring to datacenters. If you would like to keep up with what Micron is doing on this front, check out their server tech landing page.

Micron Technology’s Ryan Baxter also guested on our Utilizing CXL Podcast, part of our Utilizing Tech series. Watch his episode to hear more on Micron’s CXL efforts.


© Gestalt IT, LLC for Gestalt IT: Micron Discusses the Future of Datacenter Memory Management at OCP’s CXL Forum

View Details

Juniper recently presented a Showcase session that talked about and demonstrated two of the Paragon Automation Suite’s tools. It was quite impressive! This blog covers some highlights from the Showcase, and provides some additional coverage of one aspect of the demonstration.

But before diving into the technical side of things, let’s talk about the overall setting around Paragon Automation.

BackgroundParagon Automation is Juniper Networks’ suite of tools for automation and management of WAN transport networks (L2 and L3 EVPN, E-Line, E-LAN, and IETF models). Think MPLS and segment routing and RSVP-TE.

Paragon Automation is multi-vendor based on standards and interoperability testing. It has four main components in total that cover end to end transport network management from planning to deployment, assurance and control.

  • Paragon Planner
  • Paragon Pathfinder
  • Paragon Insights
  • Paragon Active Assurance

Showcasing The PowerParagon Pathfinder and Paragon Active Assurance were the focus of the Showcase. The presentation explained the features and capabilities, and included some nice demos.

Juniper Networks refers to the market for Paragon as the “autonomous transport network”, i.e. MPLS/SR, WAN providers, etc. In this session, Juniper emphasized on the automation and remediation capabilities of the suite in particular.

A key aspect of Paragon Automation is full observability: not only telemetry but the ability to use automated active data-plane synthetic test traffic to measure WAN network performance and SLA compliance.

Juniper has added some AI, making Paragon an automated control and remediation tool. It can automate latency-based routing, autonomous capacity optimization/congestion avoidance (re-routing traffic offloaded links while maintaining latency and other SLAs), and provide closed loop remediation (re-routing traffic away from a node or link that is unhealthy) while maintaining path diversity.

About the DemosThe demos consisted of introducing a test fault, and observing Paragon re-route the traffic to resolve the problem.

As a teaser, here’s a brief description of what you’ll see – network topology with links and stats, in the first demo case, and the latency on each segment. After a fault was introduced, the presenter showed how the path in question was re-routed by the software to avoid the problem. The discovery of nodes, links, and paths is all automated.

Here’s what the demo network diagram looked like.

The purple arc and green arc indicate to separate LSPs for the customer flows of interest.
Some other views available in the tool like link stats were also shown. In the scenario we looked at here, Paragon Pathfinder created a maintenance event on the link in question, re-routing all traffic away from said link when the fault happened while maintaining the SLA integrity of running services (in this case supported by LSPs). A prioritized approach is automated as far as selecting the first paths to re-route.

Key: Synthetic Testing!Synthetic UX (user experience) testing has been a hot topic lately, and one of the tech areas I’m keeping a close eye on.

Some aspects of Paragon’s testing capabilities are highlighted below.

For any UX testing, the key elements are the agents and tests that are available with Paragon Active Assurance. As a quick overview and reference, the following supplements what was discussed in the Showcase.

Test AgentsThe Paragon Active Assurance test agent is now natively present in Juniper ACX series routers.

The agents are also available in software-only and cloud-ready forms: VM on hypervisor, container application, software on x86 hardware, network device, on standard PC, or booted from a USB, also as securely hosted SaaS in AWS.

Test agents automatically discover and register with the control center, and appear in an inventory after being launched or connected to the network.

There is a REST API and YANG model for external control and monitoring of Paragon traffic-generating test agents.

Generally, users would deploy a mesh of test agents and tests to monitor the links and end to end services. Agents and tests might subsequently be added to add data where temporarily needed.

ControlParagon Active Assurance is operated from a cloud-ready multi-tenant Control Center with web GUI.

Types of Tests Available“Test Agent capabilities include service activation (Y.1564, MEF 48), network performance (UDP, TCP, Y.1731, TWAMP, path trace), Internet performance (HTTP, DNS), rich media (IPTV,OTT video, Netflix, VoIP telephony, and SIP), as well as support for controlling Wi-Fi interfaces, and performing remote packet inspection.” JUNIPER PARAGON ACTIVE ASSURANCE DATASHEET

The following graphic explains:

See the datasheet for much greater detail on the tests available. (Link below.)

There is also a GUI for building test sequences and specifying SLA compliance thresholds. These can be templates with parameters filled in at runtime.

What’s missing in the above list is RUM-tool like capabilities, for measuring components of an app or web page. But that makes perfect sense as it’s not something a WAN provider is likely to be doing.

ConclusionThe demos showed how easily various operational tasks might be accomplished using Pathfinder, while hinting at the automatic response the system could provide to various error conditions. They also gave a good sampling of the breadth of potential automated responses that could be enabled via the Active Assurance passive and active testing of links, paths.

That being said, looking at a couple of the demo screens and the documentation, there’s some complexity present, but it’s commensurate with the deep capabilities of the tools. At a glance, the user interface appeared to provide efficient access to those deeper capabilities.

If you’re in the WAN/MPLS/Segment Routing Transport business, make sure to check out the Showcase video, and look more deeply into the Paragon Automation Suite.

Additional links for further informationA useful repository of resources including a short explainer video, product datasheets, a complementary Appledore vendor profile and the latest EANTC interoperability test report can be found on the Juniper website. It’s a great way to learn more about how Juniper designed its Paragon Automation suite and its vendor-agnostic controller support.

You can also find the full recordings of the Showcase sessions, along with relevant case studies of Orange Poland and GARR, and summary-level demo videos, by clicking below:

  • The Autonomous Network: Key Success Factors
  • The Autonomous Network: Live Demonstration

To learn more about the topic of autonomous transport networking, any why it’s fast becoming a key differentiator by enabling superior customer experience, you’ll find free analyst reports from ACG Research and STL Partners, a CxO webinar with thought leaders from BT, Verizon and Juniper, and blog posts from Juniper executives: The benefits of end to end network automation


© Gestalt IT, LLC for Gestalt IT: Juniper Showcase: Paragon Automation Looks Impressive!

View Details

Sustainability is one of the most pressing themes of our times in IT. Every organization that is a part of this wave is pushing for it by fulfilling their obligations. At Cisco Live EMEA 2023, corporate sustainability resonated in Cisco and its partners’ commitment toward going net zero, transitioning to renewable energy, and in the overall business model. Discussion around climate change and circular economy took centerstage as executives spoke about ESG risks and opportunities, and highlighted Cisco’s goals and initiatives around environmental sustainability.

At the event, former network engineer, and Tech Field Day events lead, Tom Hollingsworth sat down with Eric Blonda, Global Alliance Executive at Cisco, and Remko Deenik, Director Systems Engineering at Pure Storage, backstage to talk about these trends, and learn about their companies’ sustainability strategies.

Turning the Page on an Era of Reckless CapitalismUntil a few decades, sustainability was not a even priority on companies’ tech spend list. Organizations operated with a downright capitalist mentality, with heavy focus on profitability. It was normal to turn a blind eye to problems like emissions and e-waste that were often dismissed as unimportant.

Fast-forward a few years, climate change sweeps through the globe, opening our eyes to the alarming reality that in order to preserve the future of the planet, industries need to reduce the impacts on environment. The first step towards that would be to embed ESG in operations, and modify business models around it.

As awareness grew, customers came at it with a growing intensity. Their purchasing habits changed overnight, making it clear that if their sustainability criteria are not met by a certain vendor, they will spend their money on a vendor that does.

“It’s the top customer concern, and we’re working in every part of our company, every part of our product line, to make sure that sustainability is a priority,” informed Cisco’s Eric Blonda.

Seeing as digital transformation is the best chance to survive in the business, companies in the last few years have started taking small steps towards sustainable practices to create more positive impacts. Today corporate sustainability initiatives are put under heavy scrutiny by authorities and customers alike.

Sustainability Pay-OffsCompanies and consumers have both come to recognize the merits of sustainability in the modern economy. Reducing energy footprint and waste inherently saves money. So, delaying the transition to a sustainable IT is not in the interest of either businesses, or their consumers.

The cost savings generated from sustainability initiatives has driven a lot of companies in the recent years, and inspired their customers to transition to sustainable practices.

Sustainability with Cisco and Pure StorageCisco and Pure Storage are jointly accelerating this transition to a circular economy. Pure Storage has sustainability embedded into their operations since many years.

“The way we’ve designed our system has been very sustainable from the start. But, the last decade, people weren’t all that interested. It’s now really picking up interest, primarily because of the energy consumption part,” said Remko Deenik.

Mr Hollingsworth pointed out, “It’s been a problem in the past in enterprise IT where we get locked into these systems where if you want to increase capacity, you have to get rid of the system you’ve been using, and you have to buy a new one because the new one is 10% faster and 20% quieter. But I could still be using the one that I was using.”

Pure Storage’s Evergreen model offers customers a break from the legacy consumption model. On it, customers can upgrade the components, and still continue to leverage the framework without overhauling it until it reaches expiration. This vastly reduces the amount of e-waste produced at datacenters, not to mention generate substantial cost savings for the operators.

“We’re upgrading all hardware components within the duration of the system. We do that disruptively, but by just upgrading the components that we need to upgrade, we minimize waste and usage of components, in addition to, of course from a sustainability perspective, shipping optimized packaging, optimized power consumption and all of that,” informed Mr. Deenik.

Green Datacenters with FlashStackRecently, Pure Storage and Cisco have launched FlashStack as-a-Service, a converged infrastructure solution that constitutes Pure FlashArray and Cisco UCS X-Series chassis, UCS Fabric interconnects and Cisco Nexus switches.

“It aligns very well, from a technical perspective, to the way we architect our system, and how Cisco’s architecting their system, using stateless design, being able to replace components at will without having to replace the entire system,” said Mr. Deenik.

Mr. Hollingsworth agreed that such flexibility would be needle-moving in the way enterprises stage their upgrade cycles, and will be able to “use the least amount of components possible to produce good performance for their users, but also good performance for our planet.”

“If you look at all the recent product announcements, you’ll see that there has always been a lot of focus on reducing complexity, and reducing the number of cables for example, reducing power, and all of that. It is part of our commitment,” said Mr. Deenik.

FlashStack uses Pure’s Evergreen for discreet scaling and Cisco’s network equipment for reduced complexity. The Pure Storage Evergreen consumption model is fully pay-per-use and lends great flexibility to consumers. Fully managed, it requires no planning at the customers’ end, and users can pay only for what they use. So customers neither overbuy capacity, nor sustain losses from underutilization

“At the start of the contract, a lot of the customers invest for the next five years. They buy a lot of empty capacity that just sits there running in the datacenter, using power for no good reason. We’re able, with this subscription model, to just right-size the solution as well as provide spare capacity so that they’re free to go wherever they want to go. We only put the equipment in place,” explained Mr. Deenik.

Additionally, the model affords flexible downsizing for times when businesses need to scale back their infrastructure when moving to public cloud. Pure Storage removes the free equipment and repurpose them to fit other customers’ environments.

Wrapping UpAchieving sustainability in an industry as robust and impactful as IT requires companies to first, believe that a clean, sustainable future is achievable, and unite in their efforts to power that future. Cisco and Pure Storage are paving the way toward environmental sustainability with their initiatives and innovations. With FlashStack, not only do the customers have a way to control cost, but also break out of legacy infrastructures to be in tune with the circular economy for better e-waste management and emission control.

To learn more about FlashStack as-a-Service, visit flashstack.com. For events from Tech Field Day Extra at Cisco Live EMEA 2023, check out the Tech Field Day website.

Panelists for Today’s Interview:Remko Deenik, Technical Director Europe at Pure Storage. Connect with Remko on LinkedIn.

Eric Blonda, Global Alliance Executive at Cisco. Connect with Eric on LinkedIn.

ModeratorTom Hollingsworth

Twitter@NetworkingNerd


© Gestalt IT, LLC for Gestalt IT: Achieving Sustainability in Datacenters with Pure Storage and Cisco

View Details

In the world of networking, there are a few widely accepted truths – managing and optimizing MPLS and Segment Routing can be problematic, outages occur at the least convenient times, and jitter, latency, bandwidth, and throughput are not terribly well understood. And yet, the potential for improving customer experience, delivering SLA guarantees, and cost efficiencies is there if you have the right capabilities.

Traffic engineering (TE) at service provider scale is a dark art, not because it entails a difficult set of tasks, but because it can be quite complex. In fact, “traffic engineering” as a term itself is vague and may have different interpretations depending on who is asked. RFC2702 describes the goal of TE as:

A major goal of Internet Traffic Engineering is to facilitate efficient and reliable network operations while simultaneously optimizing network resource utilization and traffic performance.

FoundationsTake, for example, the complexity of building a label switched path (LSP) across a geographically diverse set of devices spanning a region, a country, an ocean, or all of the above. Factor in, if you will, the requirement of a specific amount of bandwidth for that LSP.

Historically there are a few ways to accomplish this task, RSVP-TE being the venerable standby, and Segment Routing (SR-MPLS) the heir apparent. Both are high performing protocols with an impressive suite of capabilities and an equally impressive set of configurations required to fully realize their absolutely massive potential.

Configuring, and more so, maintaining these protocols can be a daunting task, and daunting tasks tend to breed innovation. Network automation has been the bellwether for massive change in how networks are built and managed, and Juniper Networks’ Paragon Pathfinder fills a unique and much needed void in automating, and thereby taming, that dark art of traffic engineering.

AlternativesAs one can surmise, there exists a variety of ways to manipulate paths in a large, complex network. Migrating traffic manually, creating time-based automation workflows using technologies such as Ansible or raw python and running them in CRON, or building a home-grown system to change the end-to-end route that a given LSP may take.

However, it should not be overlooked that leveraging the raw power of something called a path computation element, or as it is commonly known, a PCE can enable significant advantage that yields both lower operational expenditure and the potential for cost savings across potentially high cost backbone or data center interconnect circuits. By leveraging the streaming data from network nodes, real-time knowledge of the network topology, and significant compute resources, a PCE such as Paragon Pathfinder is able to create, monitor, change, optimize, and re-signal many LSPs over the topology of a large network ecosystem.

The end-to-end coverage of the PCE functionality in Juniper’s Paragon Pathfinder provides real-time, correlated awareness of network topology and conditions, allowing for the pre-computation of backup paths and actioning on path changes. It should not be understated how operationally powerful and critical path re-route is in a large, overlay based provider or data center interconnect network. It can mean the difference between millions of dollars lost, SLAs violated, unhappy customers, and productive packets seamlessly moving across another path.

Imagine now that a customer may have an application that is highly sensitive to latency and therefore, has a latency requirement. This hypothetical client has a hypothetical contract outlining the low latency requirement range. Using a PCE such as the Paragon Pathfinder, it can be as trivial as providing two end points and any other pertinent requirements, and setting the path parameters to optimize on latency, and letting the software generate the path.

Depending on the requirements, a backup path can be precalculated and readied in the case of a link failure, a fiber cut, natural disaster, or other unforeseen service interruption. And failover operates at near link-state speed, meaning any SLAs are met. This ease-of-use is available within the suite of standard protocols within SR-MPLS and RSVP-TE. So if a network is still migrating to SR-MPLS, or continues with RSVP-TE, they are able to, and if they may be a multi-vendor shop, the protocol standards should allow for a uniform operational experience.

Think now, of a capability that allows the migration of traffic over lesser loaded links. With a PCE such as Paragon Pathfinder, this becomes a realistic opportunity. By leveraging discreet engineered paths for certain traffic, it becomes possible to migrate traffic to paths that are outside of the normal routing best path scenario. More plainly, it becomes possible to easily create LSPs over longer paths in order to leverage less expensive or more lightly loaded backbone paths, thus allowing for better utilization of existing resources and saving on procurement or provisioning of new backbone paths.

ConclusionSo yes, large scale traffic engineering has a bit of the air of dark artistry, but there are ways to master and wield the power and control that it brings with relative ease. Over time, operational overhead decreases, and the potential for improved capacity utilization can increase. By utilizing a PCE any network using segment routing or RSVP-TE can centrally control and granularly control complex traffic and path engineering while at the same time avoiding the complexities of per device state tracking, and error prone manual configurations.

Additional Links for Further InformationRead up on an industry leadership perspective on this topic, with blogs, webinars and complementary analyst reports from the likes of BT, Verizon, ACG Research, STL Partners, and Juniper executives.

Juniper’s approach to end to end network automation in their Tech Field Day Showcase where you can find explainer videos, the full Showcase recording covering Juniper’s vision and perspective on key success factors, and a webinar with Orange Poland which illustrates how service providers are using Juniper’s Paragon Automation to address real-world challenges with targeted, pre-integrated use cases.

Learn exactly how this is made possible with Juniper’s Paragon Automation suite – which contains Paragon Automation.


© Gestalt IT, LLC for Gestalt IT: The Overwhelming Power of PCE with Juniper Networks

View Details

The financial industry is in a constant state of flux between moving to the cloud, and maintaining compliance with the unique requirements of on-premises datacenters. The growing need for availability and scalability, and the rising energy costs to run onsite datacenters are not helping with the back-and-forth journey.

Bare Metal as-a-Service (BMaaS), powered by Pure Storage, is offering a solution to this. By combining the economic efficiency of a consumption-based OPEX model with bare metal servers, it addresses the specific requirements of the financial industry.

Constraints of the Financial Industry in a Fast-Moving EnvironmentThe financial industry is facing a strange dilemma. Stuck between the ever-increasing demand for innovative data and availability, and the need to maneuver between private and public clouds. Organizations are struck with a new realization. As they embarked on the journey to cloud, it became quickly evident that a simple lift and shift of applications wouldn’t cut it.

An application that is not cloud-aware, or cloud-native, might not perform as expected, or not work at all. It is quite common for public cloud providers to have shared environments. Running an application in such an environment, with several tenants in a shared infrastructure, might affect performance as a tenant increases its compute requirements, or a run-away process starts to consume more CPU cycles that the provider is unaware of.

Lastly, as finance is a heavily regulated industry, the protection of sensitive data is of utmost importance. Being unable to be certain of where the data resides is a huge challenge for organizations that are already struggling with a tonnage of internal and government-issued audits and regulations.

Moving the Physical Infrastructure to a Hosted EnvironmentPure’s BMaaS is a new offering that brings a different approach to solving these challenges. It is looking to combine the benefits of cloud-based infrastructure with the autonomy of on-premises datacenters. It provides customers with full control of the hardware stack and single tenancy. And it allows institutions to configure bare metal servers to their specifications, without sharing their resources with other customers.

BMaaS is the product of a partnership between Pure Storage and Equinix which utilizes Equinix’s datacenters around the world for computing, and Pure’s data storage. All aspects of a local datacenter, from hardware to software support to lifecycle management, are under one contract and handled by one vendor.

Scalability, Data Governance, and Energy EfficiencyBMaaS allows organizations to move their datacenters and operate them in a hosted environment, built to the required specifications. Utilizing Equinix’s vast infrastructure allows firms to move their data closer to the edge, where compute and storage are needed. The same infrastructure enables data governance, as it is possible to define the region of the datacenter and comply with regulations, such as GDRP.

One of the biggest differences, when compared to public cloud offerings, is the single tenancy aspect of the service. Aside from performance issues that could be caused by other tenants, there is also an added security aspect. Meltdown and Spectre vulnerabilities surfaced years ago, and the industry is still dealing with the threat of infected systems being able to break through a hypervisor and infect other VMs running on the same host.

Another scenario is vulnerability testing on a shared system. While a tenant performs vulnerability and penetration testing against their own systems, it can have consequences for other tenants operating on the same host.

BMaaS aims to provide the flexibility to adjust to changing demands. If there is an increase in compute and storage demands, BMaaS enables the organization to spin up more resources faster than it would be possible with an on-premises data center. And while traditional data centers will occur costs regardless of utilization, BMaaS follows the cloud model as a pay-as-you-go service. Organizations only have to pay for resources being consumed.

Last but not least, an important aspect is energy footprint. As energy costs keep rising, running datacenters becomes more expensive. In recent years, one benchmark for servers that has become more relevant is performance-per-watt. Modern CPUs and higher-density servers allow for reduced power consumption. By having a hosted bare metal server, the cost for new hardware that can deliver such benchmarks shifts from the organization to the provider.

ConclusionThe financial industry has unique needs that other sectors might not be facing. With things like regulations, technical debt, and an ever-increasing demand for scalability and availability, a one size fits all approach is the least suitable solution. BMaaS presents a viable solution for organizations that want cloud like deployment and consumption benefits, but are unwilling to sacrifice performance and privacy on a shared system. It aims at combining the best of both worlds, from full control of an on-premises datacenter to the ease of availability of cloud computing.

BMaaS is available now. Learn more about it by looking at setting up a demo.


© Gestalt IT, LLC for Gestalt IT: Bare Metal as-a-Service is Packing up the On-Premises Datacenter

View Details

MPLS based networks have long supplied network engineers with a great set of protocols to allow convenient end to end service communication, while only having to configure the end nodes themselves, regardless of how many devices lie between them. In addition to configuration convenience, protocols help network engineers manage bandwidth expectations, allowing setting over-subscription levels for links, and requesting bandwidth reservations per service.

MPLS also has tools for redundancy built-in that allow automated calculation of paths and failover for node or link outages. For datacenters, WAN networks, and campus and service provider deployments, this convenience for ongoing service creation and modification is widely availed.

The automation trend also came early to MPLS network deploys for several reasons:

  • Increasingly regular new service deployments
  • More frequent topology changes via adding and removing nodes
  • General network footprint expansion

As a result, automation offers a compelling value proposition for managing these networks optimally, and reducing the dependency on manual configurations.

Juniper MPLS Automation SolutionsFrom experience, we know that home grown scripting solutions along with vendor-specific or vendor-agnostic tools exhibit a number of issues. Integrating tools, adapting to new use cases, maintaining and scaling system infrastructure are common issues faced with this approach. Conversely, for MPLS-based networks, the Paragon Pathfinder automation tool from Juniper Networks makes use of the PCEP standard to provide a vendor-agnostic tool for LSP management. Paragon Automation can optimize the planning of, and automate the creation of, LSPs on network devices, as well as expand the ability to control bandwidth and performance across the network devices.

The PCEP standard provides the mechanism for the controller like Paragon Pathfinder to dynamically create LSPs between network nodes. No configuration file change needs to occur on devices at all in this mode. Paragon Pathfinder calculates the needed LSP and communicates this to the necessary devices creating the LSP on the fly.

Controlling the LSP tunnels that transport the data point-to-point on an MPLS network, the Paragon Automation suite adds active and passive monitoring of network performance and link bandwidth utilization (not just the reservation top threshold) to dynamically assign (and re-assign) network paths.

Using the device monitoring statistics, Paragon Pathfinder knows both the maximum possible traffic from an LSP that can be supported by the underlying resources at a given time, and the actually utilized bandwidth per link across the network. As multiple LSP share links and links have an oversubscription rate, there are times when congestion can occur. For priority services the controller can detect the congestion and reroute the appropriate LSP to alternative paths in the network. Paragon Pathfinder has a holistic view of the network that can perform these calculations in a way not possible by RSVP device-based protocols.

Paragon Pathfinder process

For performance, Paragon Automation watches the latency per tunnel via monitoring. Similar to the bandwidth reroute, priority LSPs can be changed to alternative paths to route around latency issues. These capabilities are not available in standard MPLS device-based protocols at all at this time.

These capabilities are perfect for either green field deployments for new networks, or used on new services deployed in an existing MPLS network. This PCEP-based dynamic LSP creation and management works with both segment routing and standard RSVP TE based networks. But even existing brown field MPLS networks with already deployed LSPs can leverage the new bandwidth and performance features.

The PCEP standard includes a delegation option that can be added to already configured and running LSPs on existing MPLS networks. A single configuration line is added to the deployed and work LSP without any performance impact. Once in place the Paragon Pathfinder controller can monitor and change the path of the LSP. This delegation feature allows operators of existing networks to add the bandwidth and performance monitoring and rerouting to all the existing deployed services. Even the existing service deployment software, scripts or other operations can remain in place deploying new services as usual. Paragon can use Netconf or SSH from the service to deploy the delegation configuration to network devices without any operational changes needed by existing systems or scripts.

Paragon Pathfinder provides a robust PCEP controller allowing for the deploy, monitoring and dynamic control of the LSP transporting bandwidth on the MPLS network. This is also part of a suite of tools with broader capabilities in monitoring end-user service levels, providing network simulation, and network service automation platforms.

ConclusionManaging a dynamic MPLS network presents operators with multiple challenges. Making changes and adding services while also monitoring traffic levels and performance require more than just device and protocol level knowledge in the network. Centralized controllers like Paragon Pathfinder provide that system wide view and control to react as changes are required. The additional suite of products can extend the capabilities to maintenance planning simulation, service level performance monitoring and change automation platforms as well. To learn more, review the Autonomous Networks Solution description on Juniper Networks website.

Additional links for further informationJuniper has a useful web page where you can quickly access summary demo recordings of use cases for its Paragon Automation suite, including Automated Congestion Avoidance; Latency based Routing; Path Diversity. You can also find the full Showcase live demonstration that I attended recently, and links to recent case studies for Paragon Automation, including at GARR and Orange Poland. You can find all this, along with commentary from industry leaders like BT and Verizon, and complementary analyst reports from ACG Research, STL Partners, and Appledore, at the below links:

  • The benefits of end to end network automation
  • Juniper’s approach to end to end network automation
  • A deep-dive into end to end network automation use cases and case studies
  • Juniper Paragon Pathfinder: explainer video, analyst reports and key features

© Gestalt IT, LLC for Gestalt IT: Automation Tools in Managing MPLS Networks

View Details

When I was working on my expert networking certifications many years ago, I loved the idea of private VLANs. Something simple that allowed me to isolate hosts and prevent traffic leakage. Putting devices in a DMZ in a private VLAN was a great way to ensure that even if one of them happened to be compromised it wouldn’t allow those attackers to move laterally to another device in the DMZ.

Private VLANs may have fallen out of style in the past few years but the idea of host isolation hasn’t. In fact, host isolation is at the heart of technologies like Zero Trust Network Architecture (ZTNA). ZTNA is the new buzzword that solves all your security needs no matter what they might be. Unless those needs require you to have devices in remote locations. Or those devices don’t have user accounts. Or those devices aren’t powerful enough to run the entire security suite that was developed for a modern compute machine with an abundance of resources.

While we live in a world full of modern technology and more computer power than we could ever hope to use there are still areas of our lives that are ruled by low-cost CPUs. Something as simple as an ATM has a multitude of challenges when it comes to how they’re used. They may look like a computer, but they don’t act like one most of the time. They have a very specific function that requires a lot of security. You don’t want the transmission of deposit information or withdrawals to be intercepted and scraped or modified. But how can you secure a remote system in a network you don’t control and ensure the data is secure along the entire path?

VMware SD-WAN ClientI was recently a part of a Tech Field Day Showcase featuring VMware. They spoke at length about their new SD-WAN Client and how it can help you solve many of the challenges you face in a world that has become increasingly distributed. The old days of the enterprise bastion are long gone thanks to a workforce that is doing their job at home or in a coffee shop. We no longer have a castle but instead we have knights roaming the countryside. We must ensure they are armored to keep them safe.

In the video (link to published video here), Aamer Akhter talks about the VMware SD-WAN client and how it is built to function in a new distributed world. It’s more than just a piece of software for an endpoint. It’s a system of relays and connectors that allow the SD-WAN connection to use the fastest network connections to get the data to the right location. It’s a client that only initiates secure communications when needed to preserve device resources. It’s also an architecture that allows for multipath resilience and performance. Instead of hoping that you’re using the best path to the cloud or to your private data center the VMware SD-WAN Client continually assures that you are.

Going back to our ATM example, VMware SD-WAN is a perfect fit for the needs of these devices. They don’t need to send large files all the time and have no need for constant communications. Everything is transaction based and, aside from a few pictures of deposits or camera footage, are usually text-based in nature. Many ATMs run an operating system based on Windows, which would allow a client developed for that OS to work seamlessly.

VMware has taken it one step further though. Instead of using a specialized client that works only on ATMs they have built in the kinds of configurations needed for a headless client to operate on any devices. Using security tokens and device profiles you can install the VMware SD-WAN Client on the ATM and have it log in automatically without any user intervention. That means that the device will be protected even if it reboots. That should prevent wily attackers from knocking the machine offline to interrupt communications to compromise it.

Creating a client that has these kinds of capabilities is a boon to the distributed device security crowd. No longer do you have to worry about securing these devices with custom hardware or expensive solutions that require on-site technicians to configure. Instead, you just send the new ATM to the location with a tech that plugs in a hardware token and everything comes up like it should. Imagine this kind of technology working on any number of headless devices. From LED billboards to slot machines to hallway clocks the possibilities are endless. All can be secured thanks to the development efforts of VMware and their SD-WAN Client technology.

Bring It All TogetherI miss Private VLANs. Sure, they didn’t scale very well, and they were more problematic to troubleshoot. They wouldn’t work in a cloud-first environment. However, knowing I could isolate traffic was an ace in the hole when I needed it. Today’s world calls for better technology that offers similar capabilities but extends out from the local network. VMware has a real winner with their SD-WAN Client and the platform around it. Don’t limit yourself to thinking about SD-WAN as just hardware. Rethink how you want your network to operate and stay secure.

To watch VMware’s Showcase video, head to the Tech Field Day Website go to https://sase.vmware.com/sase to learn more about the offering.


© Gestalt IT, LLC for Gestalt IT: Rethinking Isolation and Security with VMware SASE

View Details

In the networked world of today, we are all consumers of some network service, and as customers, we love to be served, and most of all, heard by our providers. But seldom do we think about our service providers, which begs the questions: who serves the servers? Or which service do they consume?

Particularly in networking, especially after the paradigm shift the pandemic started, service providers work around the clock to meet the growing expectations of their customers on the services they offer.

The Paragon Automation SuiteIt has been amazing that the Tech Field Day folks are focusing on the service provider space, to talk about the challenges they have and how, many companies are serving the servers!

Some time ago, at a session with the Tech Field Day, Juniper Networks showcased their service provider focused automation offering— the Paragon Automation Suite. Though working for a service provider is definitely not a walk in the park, their offering is quite powerful, and can get even better.

Each one of the components in the Paragon Automation Suite delivers a particularly vital service that can be used by the other elements in the suite. In other words, each one of the elements uses data as the language of love within the suite, and generates an output that the other elements can ingest to deliver a service. It’s a marvelous ecosystem made of an army of foot soldiers working (and sometimes fixing/breaking) inside the computer tower case.

A convenient summary can be found in the image below:

Here are the functions that the Paragon Suite can automate:

  • Monitor the network
  • Create and deploy services
  • Modify existing infrastructure
  • Log everything
  • Save network snapshots
  • Ingest and process network data
  • Graph services and infrastructure
  • Study trends

Everything focuses on the network information. For service providers whose revenue is related to availability, satisfied customers, services, reputation and flexibility, it is a marvelous resource.

The ProblemsManaging a service provider grade network is a sizable amount of work that involves lots of technology, shrouded in the complexity that typically entails the delivery of a myriad of services to a plethora of customers.

Service providers run in similar ways as huge corporations. They have siloed teams and disjointed departments, doing more with less, performing a series of manual tasks, moving slowly, prioritizing stability and availability over anything else, and very often, preferring better the devil you know than the devil you don’t.

Rarely if ever, greenfield or single-vendor, and in many cases, definitely not state-of-the-art and novelty.

In other words:

  • Inherently complex
  • Siloed
  • Manual
  • Massive
  • Risk-averse
  • Conservative
  • Non-standard

How do you manage, operate and improve a network infrastructure of such scale and complexity? You need capable and abundant engineering resources, patience, and clever solutions. And that’s where Paragon Automation comes in!

The ProposalIn a novel package, the Paragon Automation Suite offers a series of features to address some of the biggest pain points service providers are looking to solve.

Service Modeling and Deployment

The network can be mapped and provisioned through several methods and protocols like OpenConfig and gNMI streaming telemetry, gRPC, SNMP, NETCONF/YANG, CLI, Syslog, NetFlow, and any other data sources, under the Bring Your Own (BYO) banner.

Being able to handle the word-soup of protocols above means the network devices and services from the past or the present can be added, mapped, deployed, graphed, and monitored.

Having a standardized set of models through the network allows providers to have a service-centric visibility. Knowing which services cross which devices, for how long, the sources and destinations, by abstracting the network complexity under them is helpful to say the least.

Active MeasurementOnce the service has been deployed, it is important to measure, monitor, and evaluate the trends over time. Everybody loves bars and colors drawn and to tell a story in unison, and the way to do that is through measurement.

However, without correct measuring points and mechanisms, an accurate picture cannot be drawn. If the intention is to measure what the service looks like, it is ideal to send probes from a data plane perspective. This can be done with synthetic information generated by agents in the devices and sent through the network, mimicking user traffic all the way up from Layer1 to Layer7, and looking at the services through their own eyes.

Network ObservabilityBy combining the service-centric view, topology information, network telemetry and active data plane measurements, you can get comprehensive and detailed observability throughout the stack.

Is the link congested? Do we need to move LSPs away from that device? Will it congest others if we do?

Answers to question like these can be found based on telemetry and active assurance.

SDN Solution with Multivendor SupportBy showcasing path computing capabilities, LSPs can be created, altered, rerouted, and with them, services altogether. Taking into consideration factors like shared link risk groups, affinity/coloring, priority, LSP symmetry, ERO, SLAs and path diversity, among others, services can be monitored as part of the assurance, analyzed, and their paths computed, and optimized.

That’s the entire lifecycle for the services through deployment, monitoring, and reprovisioning, if needed – all through closed loop automation, and setting the foundation for a more intelligent and self-driving network.

And the best part is, the protocols are standardized and devices can come from any vendor regardless of their already provisioned services and technologies.

Closing thoughtsThe Paragon Automation Suite offers plenty of benefits and features that could help service providers to provision, monitor, manage and optimize services. It addresses several of the common pain points providers suffer nowadays such as, manual tasks, service deployment, monitoring, activation, testing and optimization, network observability, trend analysis and multi-vendor real estate.

Paragon Automation is a rather original combination of functions and features that allows not only to supervise, but also actively influence and optimize the network on a vendor-neutral basis.

Additional links for further information:There’s a wealth of information available if you want more details. You can find everything from short explainer videos, to use case demo recordings, thought leadership webinars with executives from BT and Verizon, case study videos from Orange Poland and GARR, product datasheets, complementary analyst reports, interoperability test reports and more. Just click on the links below to get started:

  • The benefits of end to end network automation
  • Juniper’s approach to end to end network automation
  • A deep-dive into end to end network automation use cases and case studies
  • Juniper Paragon Pathfinder: explainer video, analyst reports and key features

© Gestalt IT, LLC for Gestalt IT: Paragon Automation Suite: Empowering Service Providers with Juniper Networks

View Details

When discussing network security, the approach most recommended has several iterations. “Layered Security”, “Defense in Depth”, “Zero Trust” to name a few. What these all have in common, is the idea that it is not good enough to put a firewall at the edge and call it a day. Instead, a multi-layered approach has a better chance to protect critical infrastructure, and provide staged, multifaceted security.

The New Edge Poses a ProblemWith cloud adoption becoming prevalent, the infrastructure that is being protected lives not only in a corporate datacenter, but also with cloud providers. Users and devices are far more mobile, and are expected to access resources wherever they are, at any time. The “edge” of the network is no longer a single location – it is everywhere and anywhere a user or device exists.

Traditional VPN clients initially offered a workable solution for this, by ensuring endpoints were connected back to the on-premises datacenter for access. This often means hairpinning traffic through that edge firewall, and back out for access to the internet or cloud services. This isn’t an ideal design.

The transformation of where users and endpoints work, and what they require access to, means what was traditionally defined as “the edge” has moved to every one of those users and devices, and redefined network security and protection. This has now evolved into Secure Access Service Edge, or SASE.

A Layered ArchitectureSASE remains a layered architecture. It includes SD-WAN, Firewall, Secure Web Gateway, Zero-Trust Network Access, and Cloud Access Security Broker, all working together to secure and protect endpoints and critical infrastructure, including cloud services, and traditional datacenters.

VMware discussed their SASE architecture at a recent showcase discussion with Tech Field Day. Here they offered two main options for consumption of SASE – single vendor “converged”, or two vendors “integrated”, splitting the key components of SASE into WAN Edge, the SD-WAN component, and Secure Service Edge (SSE) which is comprised of the multiple layers of technology mentioned earlier (SWG, FW, ZTNA, etc.). This provides options for organizations, those who may wish to consolidate their SASE architecture under a single vendor for simplicity, ease of management, and perhaps reduced cost, versus others who may want to handpick each of their SASE layers into a best-of-breed security onion.

What’s NewThe big news from VMware is the SD-WAN component for whichever design chosen, will now include the new VMware SD-WAN client. A VMware SD-WAN appliance makes sense for the branch or satellite office, and even the home office in cases where there may be semi-permanent, multi-device environments that need secure access.

However, for the solo user or endpoint at home, while traveling, or even working from the local coffee shop, the traditional VPN or remote access was typically needed.

This new lightweight VMware SD-WAN client is available for Windows, MacOS, Linux, iOS, and Android, and is managed via the same “single plane of glass” that all other VMware SD-WAN endpoints are managed – VMware SD-WAN Orchestrator.

VMware now offers 3 options to provide consistent, and reliable remote access for any user on any device. The traditional VMware SD-WAN Edge, the VMware Secure Access Client as part of Workspace One, and this new VMware SD-WAN Client, which combines the best of both products to include path optimization and provides access to on-premises datacenters, cloud services, and secured internet. Eventually this client and the VMware Secure Access Client will be one single product.

This new client offers the same zero-touch provisioning found in the rest of the VMware SD-WAN portfolio, including firewall and NAT traversal using the SD-WAN Client Relay. This offers fully encrypted outbound connectivity to join the SD-WAN fabric, allowing the endpoint to seamlessly communicate with the rest of the enterprise infrastructure.

ConclusionWhether an organization chooses the single vendor option, with a stacked/integrated offering that is end-to-end VMware or handpicks each individual product from a wide variety of vendors to build a best of breed security stack, each of these provide flexibility and options in how an enterprise may leverage their existing partnership with VMware. This, combined with an unprecedented number of third-party vendor integrations, gives VMware the competitive advantage to be at the head of an organizational SASE infrastructure.

To watch VMware’s Showcase video, head to the Tech Field Day Website go to https://sase.vmware.com/sase to learn more about the offering.


© Gestalt IT, LLC for Gestalt IT: VMware SASE – Architecture, Options, and a New SD-WAN Component

View Details

Secure Access Service Edge, better known as SASE brings together networking and security capabilities to better support the technology and business requirements of companies grappling with digital transformation. SASE isn’t a product per se, but an assembly of existing products like SD-WAN, remote access solutions, and security services.

The real value of SASE lies in what it delivers to an organization. SASE provides a more flexible way to secure access to both cloud and on-prem applications, and enables fine-grained security policies that can incorporate user IDs, device state, and device location.

The Two Halves of SASEAccording to Gartner, the two main components of a SASE architecture are SD-WAN and a Secure Service Edge (SSE).

SD-WAN provides core networking capabilities, dynamic path selection, and the ability to enforce business and security policies based on applications rather than ports or protocols.

SSE provides security services, which are typically hosted in Points of Presence (PoPs) in various regions around the world. The PoPs may be owned by the SSE provider, or hosted in public clouds or colocation facilities. This hosted aspect is critical to the SSE’s inherent value, that is, customers don’t have to worry about managing the infrastructure, maintaining application stacks, or updating security capabilities or signatures. They simply consume the security features as a service.

The geographic diversity of SSEs is equally important. It means that traffic doesn’t have to be backhauled to a company’s headquarters for security inspection before sending it to its actual destination. Instead, traffic can be directed to the PoP that’s closest to the user, or to the destination application directly. This design reduces the chances of introducing delay and latency, which is critical for the performance of real-time apps like voice and video.

Typical SSE security services include:

  • Firewall as a Service (FWaaS)
  • Secure Web Gateway (SWG)
  • Cloud Access Security Broker (CASB)

Other security services in an SSE may include Zero Trust Network Access (ZTNA), Data Leak/Loss Prevention (DLP) and threat intelligence.

The architectural differences come in how the SD-WAN and SSE components come together. In a converged architecture, the SD-WAN and SSE components come from a single vendor. In an integrated architecture, a customer chooses separate SD-WAN and SSE providers. Each has its benefits and drawbacks. Companies need to understand both to be able to pick the architecture that best suits their requirements.

The Convergence OptionIn a converged architecture, the customer selects a single vendor to provide the SD-WAN and SSE components. The benefits of this option are straightforward.

First, the solution should be cleanly integrated because it’s coming from a single vendor. The integration must be tested beyond slideware before signing a contract. There should be fewer GUIs to interact with – dashboards and reporting should be unified – and ideally, the converged system should make performance monitoring and troubleshooting more straightforward. For example, if a user complains about poor performance, customers won’t have to play games with separate SD-WAN and SSE providers pointing fingers at each other.

On the downside, organizations using a converged SASE architecture are limited within the security services from one vendor. These services may not be best of breed.

Additionally, these services may not align with the existing preferences, expertise, and/or processes of your network and security teams. If an organization has invested considerable time and money on a security platform that’s not included in a converged SASE, the teams may balk at being forced to change.

The Integrated OptionIn an integrated SASE architecture, an organization has to integrate an SD-WAN solution from one vendor with an SSE offering from another. For example, they could use their existing VMware SD-WAN, and then steer traffic to Zscaler PoPs for security inspection.

This option may appeal to organizations who want more choices for security services in the SSE component. This option may also align well with organizations prefering strict separation of duties: the network team can operate the SD-WAN network while the security team handles a separate SSE component.

This architecture also gives organizations more selection criteria to work with. For instance, two SSE providers might support the same security services, but one has more PoPs in relevant geographies or a more robust network backbone.

On the downside, an integrated architecture is more operationally complex. Administrators may need to stitch together various APIs to get the solutions working. There is be more than one management and reporting console.

The solution may also require more complex routing to steer traffic to separate providers. As a result, ops teams have a harder time monitoring performance and identifying the root cause of problems.

Think Architecture, Not ProductWhile making a choice, it is important to remember that it is an architectural choice, rather than a plain licensing of product. The architecture one chooses should align with the business outcomes, because it will have a direct impact on network and security operations, application performance, and troubleshooting. Vendors, such as VMware, support both converged and integrated SASE. At all times, customers must read through the fine prints to be certain that the provider/providers offer all requisite capabilities.

To watch VMware’s Showcase video, head to the Tech Field Day Website go to https://sase.vmware.com/sase to learn more about the offering.


© Gestalt IT, LLC for Gestalt IT: Choosing SASE Architecture: Convergence or Integration

View Details

Back before the pandemic, most people commuted to office daily. Then, everything changed when workers were sent home for “two weeks.” This forced remote connectivity to evolve from supporting the occasional worker to supporting most (if not all) staff. For IT Security, there was extra concern about how company data was accessed and where it was stored.

Most decent-size companies were already able to support remote access long before COVID. But this was more about granting secure access to internal systems. Plus, a lot of companies did not plan on how to support a remote workforce. With everyone working remotely, management tools like Active Directory group policies may not update like before as staff may not always connect to the VPN. Add in cloud applications, like Office 365, that staff may not need to connect to at all. Over time, it has become evident that VPN is not good enough.

Replacing VPNFortunately, there is Secure Access Service Edge, or SASE. SASE constitutes endpoint security solutions like secure web gateway (SWG), cloud access security broker (CASB), and VPN, plus replaced traditional WAN with SD-WAN.

Unlike traditional VPNs, end users should not need to think about when to connect back to the on-premises environment as the SASE solution would handle it. If end users need to access emails (such as Office 365), the connection should be direct to the cloud service. When the end user needs to connect to an on-premises server, the connection happens without the end user needing to initiate a VPN session.

Basically, there are two overall types of SASE solutions: converged and integrated. A converged solution is where a vendor includes every aspect of SASE. This makes the solution operationally simple but may not include best of breed through the entire SASE stack. On the other hand, an integrated solution is where the vendor provides the client part but connects back to another SD-WAN solution. This allows companies to use different solutions, so closer to best of breed, but having different vendor products working together could make things messier as they may not always work seamlessly.

SASE is a great tool for keeping end users safe while securing access to data, no matter if it is in an on-premises server or in the cloud. Of course, IT departments will have to transition from VPN to SASE. This could start with changing WAN connections over to SD-WAN to give a connection point for the SASE client. Then the VPN client can be replaced with the SASE client to keep that remote access functionality.

The problem is where security services have already been deployed to support remote workers. If there is already a web security product like Zscaler, does the company have to throw this out to go with a SASE provider from a networking company? Some vendors may say yes making the transition a bigger project. That is not something that may be appreciated.

VMware Can HelpVMware may not be the first to the SASE market, but they do offer an easier transition to SASE. Companies can deploy any part of VMware’s solution, or even go with the full converged solution if desired. By allowing for pieces of the solution to be installed, VMware is not forcing IT departments into a “rip and replace” scenario. Rather, the replacement can be done over time so as not to impact the business too much.

For instance, Zscaler has a secure web gateway product that processes millions of transactions daily (https://trust.zscaler.com/zscaler.net). The Zscaler client connector gets installed on laptops allowing internet access to be secured. Switching to some converged SASE solutions may require getting rid of Zscaler, which could complicate a SASE deployment. The VMware solution can allow internet access to continue through Zscaler but on-premises access goes through VMware SASE. If desired, Zscaler can be dropped later in favor of the VMware solution.

ConclusionVPNs are no more good enough to support remote and hybrid workers. By upgrading to a SASE solution, end users no longer need to make the decision when to connect over a VPN. However, replacing a VPN and existing web security solutions could be a daunting task. Fortunately, companies like VMware help by offering a SASE solution that can be deployed in stages while working with existing security services like Zscaler.

Recently, I was part of a discussion with VMware about their solution. You can listen to that conversation as part of the Tech Field Day Showcase or go to https://sase.vmware.com/sase to learn more about the offering.


© Gestalt IT, LLC for Gestalt IT: Working Anywhere with VMware

View Details

Modern flash endurance is dramatically better than early flash storage, but outdated assumptions about endurance persist. Range anxiety is causing many storage designers to invest in endurance they don’t need, instead of investing in the performance and capacity their customers actually need.

It turns out that most of the times, buyers aren’t in any danger of wearing out their flash. This should prompt a reconsideration of how endurance is taken into account when designing flash systems.

Measuring Lifetime WritesThere are two main approaches vendors use to describe drive endurance on their spec sheets: drive writes per day (DWPD) or total bytes written (usually as terabytes written (TBW)). They’re subtly different.

DWPD doesn’t really match the way users think about writes when using a system. They tend to think about throughput, which is about total bytes. At a given throughput rate, how long would the system last? That’s easier to work out if thought in terms of total bytes written.

A 7.68 TB flash drive with a rating of 1 DWPD would be able to handle 14,016 TB of writes before its industry standard 5-year warranty expires. That means for 20,000 TB of endurance, the drive needs to be bigger, right?

Not necessarily. Some 7.68 TB drives are rated differently for random vs. sequential writes. This is because sequential writes are more predictable, and the drive firmware can better manage the way writes are performed, extending the life of the drive. Sequential-write optimized 7.68 TB drives can support over 20,000 TB of writes in 5 years.

Costly mistakes can be sidestepped by bringing vendor drives specs into a common language, like total lifetime bytes. And if the kind of workload is matched to a drive optimized for that workload, then better investment choices can be made.

In fact, users probably don’t even need the endurance they think they need.

Actual EnduranceA large study of drive endurance found that most drives had used less than 15% of their predicted lifetime endurance. Drives simply don’t get used enough to wear out in most enterprise installations.

It turns out that drive firmware appears to be the most significant factor affecting drive reliability. Bugs in the software are more likely to trigger problems that cause drives to be replaced, not overuse of the drives that wears them out early.

Familiarity with workloads and drive behavior, which is used to make firmware improvements, is therefore something to keep in mind when selecting a flash drive vendor.

Given the discussion of drive endurance specs above, focusing merely on DWPD or total bytes written without an understanding of workloads would be a mistake.

p.p1 {margin: 0.0px 0.0px 8.4px 0.0px; font: 9.0px Helvetica; color: #000000} span.s1 {text-decoration: underline ; color: #0000ff}
Captioned: “Annual replacement rates for different drive families, grouped by firmware version. (Source: Stathis Maneas et al., “A Study of SSD Reliability in Large Scale Enterprise Storage Deployments,” 2020, 137–49, https://www.usenix.org/conference/fast20/presentation/maneas)”A Word on PE CyclesThe rules of thumb used to be that each increase in cell-level technology (single-, to multi-, to triple-) dropped the number of PE cycles by about a factor of 10. Thus, with starting SLC cells starting at around 100,000 lifetime PE cycles, TLC cells would end up with around 1,000 PE cycles.

This log-linear approach is easy to remember, but it’s not quite accurate. Enterprise MLC cells were able to get 3 times the PE cycles of regular MLC (30,000 instead of 10,000), and TLC cells also managed 3x the log-linear predicted cycles (3,000 instead of 1,000).

When the discussion gets to QLC 3D NAND, things are very different. Instead of the rule-of-thumb predicted 100 PE cycles, one can get at least 1,000. In fact, modern QLC 3D NAND drives with their vastly more sophisticated firmware are regularly able to achieve 2,000 or even 3,000 PE cycles, which brings them into TLC level endurance range.

Naïvely assuming that there’s a substantial difference in PE cycles between modern QLC flash and TLC flash is going to lead to wrong decisions. It’d be foolish to rest on assumptions and not look at what modern drives are actually capable of.

ConclusionModern enterprise flash lasts much longer in real-life situations than users tend to assume. Drive reliability has very little to do with theoretical endurance under most enterprise conditions.

It’s time to get over range anxiety and start looking at the wealth of real data that modern storage systems offers. If the industry can let go of outdated fears, buyers may well be able to invest in storage systems that provide superior result for them.


© Gestalt IT, LLC for Gestalt IT: Modern Flash Endurance with Solidigm

View Details

One of the advantages of Kubernetes development and operations is the modular construction of the infrastructure itself. This advantage is extremely prevalent when it comes to integrating storage: Kubernetes has the right level of abstraction so it can keep a loosely-coupled connection to the back-end storage subsystem. Although the Kubernetes Container Storage Interface (CSI) is a helpful abstraction, one long-time challenge is that it limits the number of platform-specific features that can be utilized by containers. This makes Dell’s new Container Storage Modules (CSM) approach a huge benefit, since it unlocks all the features at the hardware and software storage subsystem to bring those capabilities closer to the applications running on Kubernetes.

Abstractions With AdvantagesContainer Storage Modules (CSM) are built on top of the Container Storage Interface (CSI). This allows abstraction to remain while also getting more intelligence from the container platform layer, bringing the benefits of scalable enterprise storage platforms closer to the cloud-native nature of Kubernetes. CSM removes the need to program key features into Kubernetes itself, reducing complexity in the core Kubernetes codebase.

The CSM concept creates appropriately-opinionated modules that are loosely coupled. This is the ideal level of interoperability, because any underlying storage system with enhanced capabilities can expose those features using CSM as a common method.

Who Will Use CSM Features?The primary consumer and administrator of storage in traditional enterprise environments is a storage administrator (or perhaps the operations team). But companies are now creating a new practice group, dubbed platform operations. Even though they are more developer-centric, the platform ops teams are still on the “Ops” side of DevOps.

CSM now allows developers to pick and choose how to leverage features since it allows programmatic access to them. The current modules that are available and fully supported by Dell today include authorization, replication, resiliency, observability, and volume snapshotting. There are more in tech preview today (app mobility and secure), and we expect that these will reach production-level support soon.

It’s easier to think of the value if we look at some active use-cases that Kubernetes operators and cloud-native application developers will know very well.

Use-Case #1 – Scaling Container StorageThe original intent of containers was ephemeral computing, but this has shifted. Today, long-running workloads and shared container storage have become common as DevOps teams seek to get the most out of Kubernetes as a hosting platform without having to refactor their applications to be 100% stateless.

The issue with stateful, long-running workloads is that they can have operational patterns that aren’t “typical” for Kubernetes. One of the challenges is how to scale storage without re-spawning containers to a new location. Application developers and operations teams have been holding back app migrations to Kubernetes because of the limits to storage scalability.

Using CSM to abstract the underlying storage lets developers present and manage the storage endpoint programmatically. They will not have to worry about requesting access through a ticketing system to make changes because they can either be given access to manage it themselves or the operations teams can use simple, programmatic methods to operate and expand, contract, or set properties like tiers and capabilities for the storage system.

Use-Case #2 – Authorization For Container StorageAccess to storage is traditionally managed at the cluster layer, and there has not been effective granular access based in the core RBAC for Kubernetes. This is very risky and leads to issues like allowing unnecessary access or, even worse, leaking authentication code and secrets.

The authorization module in CSM remedies this concern through the programmatic selection of a storage target for the container based on some criteria (e.g. location, encryption state, type of storage). It also allows for programmatic management of authorization.

CSM authorization allows application developers to include storage access in their process with much less risk exposure. It also makes quotas and other properties of storage management much easier for the operations teams without having to manually allocate at the storage subsystem for every request. Authorization is now exposed to the container with the native CSI and existing properties that are already part of storage management in Kubernetes.

Use-Case #3 – Cross-Cluster Data Replication and RecoveryAnother common challenge facing Kubernetes DevOps teams is cross-cluster data replication. Application failover and recovery to alternate clusters is a challenge when complex data and applications are involved. CSM allows for snapshots to be created and replicated to a secondary cluster. When the new containerized application is spawned in that cluster, it will have the latest revision of the data. This array-based replication for cloud-native environments allows DevOps teams to realize higher performance, simplified management, and an efficient use of resources. In particular to performance, the replication available through Dell’s CSM helps enterprises avoid the overhead of data transfer being done in software above the storage hardware, instead completing the replication at the storage array level.

The unique advantage is that CSM methods can combine all three of these use-cases. By having scalable underlying storage that does not require container and node restarts to recognize configuration chances is game-changing. The ability to authorize storage access extends quota management and storage type access management across the whole environment instead of just inside each cluster.

This is a fantastic solution for business continuity and disaster recovery and gets rid of the need for the replication to be handled inside the application itself. Considering how developers manage storage today, they will be very excited to hear about using CSM for Replication.

ConclusionKubernetes and cloud-native application design is a big change for many organizations. The lack of specific performance-oriented support of powerful underlying storage systems has long held back application modernization. CSM brings features like replication, resiliency, extended observability, and volume-level snapshotting to ensure programmatic control and the right abstraction to keep Kubernetes core as simple as possible.

The work being put out by the Dell platform team to expose and fully leverage storage features using Container Storage Modules is a huge win for enterprises adopting Kubernetes as a containerized application hosting platform. The team has produced lots of helpful CSM resources, from the core guide to specific live GitHub content for each of the associated modules. CSM is the beginning of a new era in storage integration that will hopefully drive more innovation by the entire Kubernetes community and supporting vendor ecosystem.


© Gestalt IT, LLC for Gestalt IT: Kubernetes Storage Gets Turbocharged with Container Storage Modules with Dell Technologies

View Details

For decades we have seen incredible innovations in compute, storage, and networking. Virtualization, public cloud, and more recent adoption of cloud-native technologies like Kubernetes have become the new landing spot for many modern applications. Multi-cloud deployments are becoming a standard pattern. It can be by design, or simply because many applications have deployment needs that tie them to different cloud providers. But all of these innovations still leave a consistent challenge for developers and operators alike: How do we protect our applications and data as the infrastructure patterns change?

Data Protection is BrokenData protection technology works well, but the implementation is often broken in modern applications. There has been a lack of focus on just how complex and costly data protection can be, especially in multi-cloud and hybrid cloud implementations. Traditional approaches to data protection involved layering software on top of applications and file shares and then moving copies of data to another storage destination. Even in traditional datacenters this approach loses any efficiencies achieved by the primary storage solution since the orchestration and data movement is performed outside of where the data is stored.

Applications are becoming more distributed, disaggregated, and complex. There are clear application-related advantages gained from using microservices and distributed application design patterns. The complexity tradeoff comes along with designing how to protect and recover those applications. That complexity leads to increases in both operational and people costs.

Multi-cloud amplifies the complexity but is virtually unavoidable as companies take advantage of the best capabilities of each cloud. The result is that DevOps teams have to build data protection processes that fit each cloud provider.

Complexity Comes by Design Modern applications are generally being built with multiple application services, data services, and distributed architectures that also connect to other applications. It is complex enough to protect each tier of the application, but is exponentially more difficult in complex multi-cloud environments.

Multi-cloud complexity comes from multiple sources, including the following:

  • Authentication and Authorization – Identity and access management is vastly different between the major public and private cloud providers. This means that backing up and restoring the application requires unique identity and access to be a part of the process.
  • Proprietary APIs and Services – There is no “one size fits all” API to each cloud, or services across the clouds. Compute, storage, and networking will behave and cost differently between different clouds.
  • Cost Management – Each cloud has its own cost model, reserved capacity model, and also contract-related bulk purchasing and discount options for enterprise customers. This makes it very difficult to predict and manage the costs for primary applications and their data protection requirements.
  • Consistent Change by the Provider – Public cloud platforms innovate at a breakneck pace which means having to keep up with rapidly changing applications and infrastructure in your data protection designs.
  • Inconsistent Data and Storage Service Models – Each cloud has proprietary database services, key-value stores, object storage, block storage, each of which have different service tiers and pricing.

It’s easy to see how this adds complexity to your data protection strategy. Operationalizing that strategy usually leads to even more complexity in practice.

The Cost of Multi-Cloud ComplexityA good data protection strategy should include in-site and off-site. When a “site” is a public cloud provider, we have to account for storing and recovering data in different regions. This requires more compute, networking, and access management in that cloud.

Many teams are also looking at how to store data from Cloud A into Cloud B for resiliency but it’s unlikely that an application can be recovered in Cloud B without a great deal of work. Another hidden cost that comes in real dollars is data transfer charges across regions and between clouds. Primary, secondary, and long-term storage is also complex to estimate pricing for. It’s likely that a third-party solution will be needed due to the continuous complexity battle in a multi-cloud design.

Kubernetes Komplexity Kubertnetes is quickly becoming the most popular multi-cloud platform. It has become a common underlay regardless of which cloud or on-premises provider is used. This is a huge win to reduce complexity for the applications but it also comes with its own tradeoffs.

Data management and protection in Kubernetes adds a whole new layer of complexity. While containerized applications may be stateless, the data that they read/write/modify is probably stateful and needs to be persistent and centrally managed.

The applications and processes traditionally used to protect IaaS and VM-based workloads won’t work for containerized workloads. That puts applications at increased risk and adds complexity for operations teams.

Data Protection is Actually Application ProtectionWe talk about data protection all the time but the data is really there to support applications: Business may run on data but that data is managed and used by applications. Businesses must align the application with the data, and build data protection to map to recovery objectives.

The RTO (Recovery Time Objective) and RPO (Recovery Point Objective) will be individual for each application. Data protection has to align operational processes with policies and requirements. For example, an organization must size the media or backup servers appropriately to stream and process the backup data. But, this can come at a cost to the production storage which must treat the protection application as an additional workload.

The following important questions about data protection must be considered for each application:

  • What services and components make up the business application?
  • What cloud dependencies does the application have?
  • What data does the application need?
  • Which applications are also sharing this data?
  • What are the RTO and RPO of the application?
  • What is needed to protect and restore this data based on the application recovery requirements?
  • What is the plan for creating immutable storage options and backups in case of data issues (e.g. ransomware, viruses)?

This is the base from which we can define a data protection strategy, and it will influence how to build and manage a primary storage solution.

NetApp approaches the problem differently, since their storage is designed to efficiently create and quickly move data copies. When it comes to data protection NetApp ONTAP can effectively eliminate backup windows. Rather than using file-based streaming, ONTAP data protection creates space efficient copies of data using snapshot technology. ONTAP then moves the snapshot copies to secondary storage or object storage, using block-based data replication, in 4k byte chunks. The advantage is that only the changed bits move, and not the complete file. Another benefit to this approach is that all of the storage efficiencies of data deduplication and compression are retained from source to destination. The result is far less data movement, which accelerates protection and minimizes load on the production storage. Recovery time as well as data loss can be dramatically reduced as more frequent copies can be made and stored. All of this can be orchestrated out of the data path by BlueXP.

Additionally, placing the responsibility for protecting data on the storage means that there is less room for error and complexity. The storage system is aware of the location and distribution of data copies, resulting in one less point of management. As the application environment becomes more dispersed, deploying protection with the data is a more efficient and secure approach compared to more complicated streaming solutions.

ConclusionThere will be more complexity and costs as companies build data protection for multi-cloud environments. It’s a matter of managing the tradeoffs using data storage and management platforms that can provide consistency across multiple clouds.

The more consistency we can create, the easier it will be to protect and recover the applications. Multi-cloud data protection is not simple, but it’s necessary. It’s got to be a goal as operators to find the ideal platforms and solutions to meet the needs of applications and create operational consistency wherever possible.


© Gestalt IT, LLC for Gestalt IT: The Multi-Cloud Data Protection Conundrum

View Details

Housing and backing up Microsoft SQL data has always had its challenges. Database administrators want control of their data, and they want it close to the processing. Storage admins have to contend with resolving data gravity issues while finding ways to back up the databases across different platforms and protocols. With Microsoft announcing new storage integrations with SQL Server 2022 at the recent PASS Data Community Summit, and using the power of the Pure Storage FlashBlade//S, DBAs and storage admins can finally find an easier path forward to these problems.

S3 and SQL Server 2022At the 2022 PASS Summit, Microsoft announced that the next release of SQL Server would feature S3 (AWS Simple Storage Service) Object Storage integration. This is going to offer up Parquet, CSV, and Delta file support with the initial release. The SQL Server 2022 (16.x) release’s storage offerings are going to be more robust right out of the box, as opposed to gradually releasing additional file types. This could not have come at a better time considering Microsoft’s announcement that SQL Server 2019’s Big Data Clusters would be depreciated in future releases, impacting the ability to scale.

Pure Storage and Microsoft’s RelationshipThere is a strong relationship between Pure Storage and Microsoft, including many Microsoft MVPs on Pure’s staff. With this relationship, Pure Storage has assisted Microsoft in the research and development of the S3 integrations with SQL Server 2022. While not the only partner in the ecosystem to help with this development, Pure Storage has been a leading alliance in this project.

Where Do Pure Storage and FlashBlade//S Fit?Historically speaking, backing up SQL data away from the original storage array was a multi-step process when using S3. Pre-S3 availability, the SQL databases that ran on Pure Storage FlashArray, the data would be backed up to another FlashArray, and then a 3rd party tool would be utilized to convert to S3 and migrate to a FlashBlade. With the new announcement, the entire process is streamlined and data is moved directly to the FlashBlade, simplifying the work of both DBAs and storage administrators.

Besides simplifying the process of protecting the SQL Server data, the addition of S3 and the use of Pure Storage FlashBlade//S has certain benefits. The processing and storage horsepower of the //S allows storage administrators to scale out their SQL environments, rather than just scaling up.

Pure Storage also allows for a single namespace to be utilized. With the power of the FlashBlade//S’ multi-threading and read/write buffers, enables accelerated backup and restores. Data provided by Pure Storage also indicates that as the number of databases on an array grows, the speed at which data is backed up and restored also increases.

With the S3 protocol utilizing REST API, the syntax to configure backup and restoration targets is similar to what they already know. Pure Storage anticipates no barrier-to-entry or learning curve with the release of S3 for SQL Server 2022.

Final ThoughtsTime and time again, the innovation and collaboration that emerge from the team at Pure Storage elevate not only their offerings, but also what the partner ecosystem can deliver. By working with Microsoft, Pure Storage and others have simplified the day-to-day operations within the datacenter. DBAs and storage administrators can rest well knowing that the process to protect the SQL data has been improved.


© Gestalt IT, LLC for Gestalt IT: Improving the Relationship Between Storage Admins and DBAs with Pure Storage and Microsoft

View Details

Data is one of the most important assets for businesses around the world today. It lives in a lot of different places – on premises, in a private cloud, in a public cloud, or in hybrid or multicloud. It is created and moved every day. A lot of this data is created as files. Storing these files in different locations while still maintaining the desired control and security can be challenging.

To mitigate that, NetApp delivers a seamless experience across on-premises and public cloud environments with NetApp BlueXP – a unified control plane that comprises multiple storage and data services delivered via a single SaaS-delivered multicloud control plane.

With BlueXP, NetApp offers customers not only the ability to utilize Cloud Volumes ONTAP, but also services on public clouds like AWS FSx for NetApp ONTAP, Azure NetApp Files, and NetApp Cloud Volume Services for Google Cloud. BlueXP also supports a host of other data services, such as observability, governance, data mobility, tiering, backup and recovery, edge caching, and operational health monitoring. Last, but certainly not least it supports ONTAP 9.10 and beyond deployments in all public clouds in addition to on-premises deployments. But before we first dive deeper into the future, let’s first look at the past.

The Present of FilesThe NetApp ONTAP operating system embraces traditional, SDS, and cloud paradigms. It provides customers with the flexibility to deploy ONTAP for different workloads on architectures: on-premises on optimized AFF and FAS appliances, or in any major cloud as either a self-managed or a fully managed offering. Customers get a unified view across all data and assets, consistent policy application, and simplicity of management. Cloud Volumes is implemented as a global namespace that abstracts multiple deployments and locations regardless of distance.

Based on ONTAP, Cloud Volumes is architected to support hybrid deployments natively, whether on-premises or in the cloud. Tiering, replication, and data mobility capabilities with NetApp have always been outstanding. Topping that, it enables an amazing experience that lets organizations decide where data resides, how and when it gets tiered, and where data copies and backups used for disaster recovery should be replicated to.

The Future of FilesAll these operations can be directly executed from the BlueXP management interface without having to access each public cloud console, drastically reducing the time spent on usually tedious operations. The future of File-based storage will be greatly improved by data management solutions like BlueXP. NetApp always had strong data analytics capabilities through which data management can be done through BlueXP which has integrated dashboards with NetApp’s Cloud Insights service), and through Cloud Data Sense, now referred to as classification as part as BlueXP. This provides insights around data owners, location, access frequency, and data privileges, as well as potential access vulnerabilities, with manual or automated policy-based actions. Organizations can generate compliance and audit reports such as GDPR, HIPAA, and more. Other regulatory reports also can be run in real time on all Cloud Volumes data stores.

Security is the next “future” service in which BlueXP provides advanced security measures against ransomware and suspicious user or file activities when combined with the native security features of ONTAP storage. The Ransomware Protection dashboard, available in BlueXP, monitors security and user behavior to help identify risks and threats and instruct how to improve an organization’s security posture and remediate attacks.

ConclusionNetApp keeps innovating and bringing services to organizations that help managing the ever-growing data. The future of File storage has been in NetApp’s DNA for decades and copying this into the “cloud” solutions means that file storage is, and will continue to be a very flexible, secure, and capable solution for organizations of all sizes.

For more on this topic, take a look at the Take Charge of your Multicloud Environment with NetApp event on the Tech Field Day website, where you’ll find more videos on this.


© Gestalt IT, LLC for Gestalt IT: File Storage in the Multicloud with NetApp

View Details

Tamper-evident technologies have existed for decades on food and drug packages. We see them every time we open a new bottle of over-the-counter pills, or crack the bond of that ring of plastic attached to the milk jug lid. They are there to assure us that the product inside has not been tampered with or altered. They are there to prove the integrity of the thing we bought.

Technological IntegrityIT supply chains are also important to ensure integrity of the products we utilize inside our datacenters. This integrity is part of the CIA security triad (confidentiality, integrity, and availability) that can be used to secure data. Determining the integrity of the system should include verifying the system is exactly what the manufacturer intended it to be, without manipulation.

For items like Internet of Things (IoT) and Operational Technology (OT) devices, the integrity of the product is generally the combination of the hardware, OS, and applications, which rarely changes and is defined during manufacturing. Nonetheless, proving the integrity of these systems can be complex due to the variety of silicon, manufacturers, and ownership transfers that exist in the build and supply chain of these devices.

Verifying Integrity with Micron AuthentaThis is the challenge that Micron’s Authenta platform was designed to solve. With this technology, integrity can be verified through a combination of technology embedded in the lowest level of the system and an external certificate that can be combined to provide verification of the system.

The first part of the solution is the Authenta Root of Trust. This is a secure element that is embedded into standard memory products. Each element that Micron manufactures has a digital key that can attest to the hardware configuration throughout the manufacturing process. Utilizing this key, the OEM can create a certificate that can ensure the storage on the device contains a known-good image and unaltered hardware configuration prior to booting. Device integrity can also be validated while the device is booted and in use. Even more powerful, since the element is built into the data path of the system, it can actively block attempts to modify protected sections of storage.

The second part of the solution is the Authenta Cloud. With Micron’s manufacturing capabilities, they can easily build the trust in the silicon, but there still needs to be a secure and scalable way to share the keys and store the device certificates across the lifecycle of the device. For that, Micron created and operates a cloud service that provides a cloud-based key management system (KMS) and certificate authority. Utilizing this system, any trusted partner in the supply chain can validate the device they’ve received has not been tampered with and update the certificate that can be used to verify the system they manufacture.

When creating the Authenta technology, Micron didn’t set out to force OEMs to use a specific system in order to participate. With heavy emphasis on APIs, the Authenta system can be utilized within the OEM’s existing manufacturing systems. This includes the ability for the key and certificate to be transferred through a closed, pre-existing, or non-cloud system. The most important element is that they are transferred independent of the device itself, thus creating a bifurcated supply chain that can be used to validate each other. With this model, both branches would need to be compromised in order to create an undetected supply chain attack.

More Than Just a Safety SealIn the end, we can imagine a security camera or smart lightbulb that can be constantly checked for manipulation and verified to not contain any unintended hardware or software. With a “golden config” stored and checked in the cloud, OEMs could even deploy updates and new features over-the-air automatically. With Authenta running within the flash device, malicious writes can be completely prevented or overwritten with the golden config on the next boot.

This ability to provide automatic, silicon-level integrity can be an important element of a zero trust architecture. It provides zero trust in the device before it boots up, with Authenta providing the ability to verify it is a trustworthy configuration before it boots up.

While it’s still important to spend time ensuring that users are buying from someone they trust, Authenta does allow them to know they are getting something with the integrity of the manufacturer. With Authenta, users can check that safety seal and know it’s exactly what the manufacturer intended it to be.


© Gestalt IT, LLC for Gestalt IT: Micron Authenta: A Safety Seal for the Digital Age

View Details

As companies have progressed into public cloud, they have faced many challenges starting with differences in security and cost overruns due to misconfigurations to simply having the staff in place to manage complex cloud deployments. To add to this, many organizations have evolved into using multiple clouds for their workloads—whether that is because of mergers and acquisitions, plans for high availability, or shadow IT adopting technologies without the knowledge of the IT org.

Beyond multi-cloud, another challenge IT groups face is storage performance. In most cases, native cloud storage performance may not meet the demands of mission critical applications like SAP HANA, or other business critical cloud systems. Having all these new, distributed computing platforms means it can be extremely challenging to understand where your data is, who should and does have access to it, and how you are protecting that data. While these problems are challenging, there are ways to overcome them.

Complexity of Managing Multi-Cloud EnvironmentsWhile sometimes firms choose a multi-cloud strategy for regulatory or availability reasons, in many cases, the decision is not made by the central IT organization. However, those same organizations need to manage the cloud infrastructure to protect their compute and data assets. The challenges of supporting disparate cloud platforms is mostly a people problem. While the concepts are similar, implementation details around security, networking, and storage can be quite different between each provider. This means in a lot of cases, organizations will have a dedicated team for Azure, AWS, and Google Cloud, and work shuffled to each of those teams.

One of the ways to reduce the burden of supporting those differing platforms, is to use abstraction layers. For example, one of the reasons why Kubernetes has become a deployment platform of choice for DevOps teams, is that you can deploy from the developer workstation to three different clouds, with only minor changes between a deployment script. Likewise, with storage, using a common platform like NetApp Cloud Volumes can simplify storage management by supplying a common storage platform across clouds. By offering a common storage platform across clouds, NetApp helps eliminate storage complexity between clouds.

Storage Performance ChallengesOne common challenge faced by most organizations is the lack of storage performance in the public cloud. While the cloud providers have improved their bandwidth and latency numbers over time – since most cloud storage is remote (and object-based) – it can be difficult to match on-premises performance. Yes, there are some VM types that support considerable volumes of local NVME storage, but that storage is ephemeral (meaning that if the VM is deallocated during a host failure, the data on those disks are lost). It requires an organization to build their own redundancy model on top of that storage, which is not practical for most applications.

For applications that are dependent on low latency, high throughput storage like SAP HANA (which has its own hardware certification matrix), or mission-critical enterprise databases running on SQL Server or Oracle, these challenges with cloud storage can be hard to overcome. While it is possible to build low latency storage in the public cloud, these configurations can add more complexity to the architecture, and require specialized skills to support. Having a simple, high performance storage solution can deliver the performance applications demand without overwhelming complexity. NetApp Cloud Volumes can deliver high levels of throughput and IOPs, at single-digit to sub millisecond latency.

Data GovernanceAs workloads spread across multiple clouds, it can be hard to keep track of what data is where. Data governance is not the most exciting topic, but it is critical both in terms of ensuring that you know your sources of truth from your data and supplying data security. For most, security may not be top of mind when talking about governance, but having a clear map of where your data is stored, and who has what access to that data provides a roadmap to how your data should be secured, and in the event of a data breach, understanding exactly what data left your office can limit the scope of GDPR fines.

Implementing data governance policies can be a challenge but using a storage platform that supports easy-to-use labeling, encryption at rest, and supplies automation and reporting can make it much easier to meet the regulatory requirements. NetApp’s Cloud Data Sense can help discover and profile data to provide enhanced governance and privacy.

ConclusionMoving into the cloud is a big challenge for most organizations. Trying to extend that migration to a multi-cloud solution can be incredibly challenging for even the best IT organizations. Business demands can force IT organizations into those multi-cloud scenarios, despite those technology concerns. Having abstraction layers that can supply performance, easier management, and deliver strong governance practices can help counter and overcome cloud challenges.


© Gestalt IT, LLC for Gestalt IT: The Challenges of Multi-Cloud: Storage and Beyond

View Details

At the end of every fiscal year, management takes inventory of current assets and dictates the goal for the upcoming year. This determines what future budgets will be specifically built around. Every organization knows these to be annoying and tedious tasks. Modern practices such as DevOps, and FinOps are built around helping these tasks be more predictable, streamlined, and easier.

Spreadsheets like the one above, are the lifeline of every company, but they get out of date quickly. They also don’t dynamically show what resources organizations are consuming.

Cloud CostsWith the adoption of cloud, understanding how to quantify each year has become difficult. Where the regular four-year enterprise license agreements (ELA), and end of life (EOL) purchases were easy to see coming, cloud made these costs smaller, more dynamic.

Unfortunately, with scalability and elasticity comes complexity. On-Demand resources are easy to get started in the cloud, but they are costly and difficult to predict. Spot instances and reserved instances grant nearly fifty percent reduction in cost but are only good if the workload used fits the infrastructure deployed.

Price is Predicated on PrescriptionFinOps is the methodology of finding financial means to fulfill operational requirements, and while it is a great model, the tools are the critical component. Till today, few accountants know the difference between EC2 and Lambda in AWS public cloud. This is where tools are required to help organizations answer critical operational, and financial questions, such as, how to reduce costs without sacrificing performance, or to keep from building unnecessary projects or resources. This is where tools like NetApp BlueXP, Cloud Insights, Cloud Tiering, Spot by NetApp, and other solutions come in to help users quantify what they have, what they need, and most importantly, what they don’t need.

NetApp Cloud Insights and Cloud Tiering focus on optimization and observability of resources and that is a key differentiator. Before organizations can discover what they should consume next year, they need to know how to validate what they are currently consuming, and how to optimize it. These are key steps prior to implementing the full suite of FinOps. In addition to these tools, FinOps requires something like NetApp BlueXP which gives organizations options for resource management based off cost. This combines observability and optimization with predictable financial planning.

NetApp BlueXp Digital WalletNetApp BlueXP Digital Wallet uses a “bucket” type financial model that allows users to understand what they are spending, and what their resources cost, starting with services and licenses they are using in the cloud and on-premises. These licenses help keep a tab on what is expired, what is being used, and even exchange an unused license for something else. This is fantasticbecause most FinOps tools do an excellent job at observability and notifications, but the ability to manage these cost centers through a self-service portal is an excellent addition.

Licenses being swappable in a quick, easy, and self-service way is the most cloud centric service seen in FinOps and is a great differentiator with BlueXP. There are vendor licenses out there that are a group of individual licenses and cannot be swapped for the individual licenses. Enterprises see this as simply the cost of business, and not a broken system. BlueXP engages operations and financial teams, allowing them to set what they need, when they need it. The technology shows where we may see FinOps moving to in the future.

ConclusionFinOps as a practice is still young. This is a new perspective, because cloud engineers are tired of trying to figure out what is costing so much, and where. They want to be able to build, or not build based off the projects they are assigned. They do not want to have to spend hours sending financial data to their managers. NetApp is ahead of the game, not only giving visibility, optimization, and automation to help users see what they are doing and what it costs, but also granting self-service functionality to help configure the cost, and the licenses. That’s what makes NetApp’s suite so engaging and exciting. While FinOps will grow from here, I’m excited to see what the future holds.


© Gestalt IT, LLC for Gestalt IT: A FinOps Perspective: Falling out the Manual Trap and Getting to Predictive

View Details

It’s safe to say that we’ve reached the point of cloud adoption where no one is asking what containers and Kubernetes are. The ubiquitous nature of these platforms has shifted the conversation away from on-premises data center deployments and firmly entrenched the discussions in public cloud for the foreseeable future. Enterprises that are looking to the cloud are planning for Kubernetes in a big way. Net-new application development and deployment are container-focused. Existing virtual machine (VM) deployments are being refreshed as well but they are also receiving attention as containers allow them to be augmented and extended in new ways.

Moving to ContainersKubernetes smooths the transition between writing legacy code and creating cloud applications. Instead of worrying about the specifics of how a given environment operates you can just write your application to use Kubernetes and be sure that it will run on any flavor of cloud. This kind of portability makes the decision makers in your organization happy because it means they can move without worrying about anything breaking along the way. Sure, there’s a lot of consider when writing that code but knowing you could redeploy from Cloud A to Cloud G whenever you wanted is a very appealing thought.

Does that mean that containers are the answer to all our infrastructure woes? Does the nature of microservices development mean we can finally do away with protecting the infrastructure stack and just wing it? No doubt you’re already shaking your head in disagreement. No matter how transient and abstracted the layers are we still have important information that needs to be protected. Whether it’s the critical data that our organization creates or relies upon for decisions or the way our infrastructure is created there’s far too many things that can cause problems should they just disappear.

When you look back at the history of Kubernetes adoption you don’t see the all-in mentality of deployment. No enterprise woke up one day and decided to shift everything to containers. Instead, they took a measured approach. They deployed non-critical applications using containers and worked through the bugs. Enterprises are very risk averse and they don’t want to risk going out of business because of a bad bet. IT knowledge workers are also risk averse in the enterprise, since one major issue could lead to them creating a new resume container and deploying it on the job market.

Like any new technology, containers and Kubernetes needed to mature before the enterprise was ready to embrace it. Sure, DevOps loved to deploy things on this new platform. However, the mantra of moving fast and fixing broken things along the way doesn’t resonate with the decision makers. They’re not ready to make the move until they can be sure that this new bet isn’t going to sink the company and their stock price. How can decision makers feel more comfortable about joining exciting new tech to a risk averse mentality?

Commvault Means ProtectionWhen you think of Commvault you think of protection. You probably think of a program that backs up your user data and stores it somewhere safe in case disaster strikes. You may even know them for their new focus on ransomware protection. But do you think of them when you think of Kubernetes? Did you even know they worked with it?

Commvault can back up your Kubernetes workloads. Not only can they protect individual workloads but they can protect your namespaces as well. Maybe you want to protect your production environment without backing up the chaos that exists in your development namespace. Commvault has the tools to make sure only the namespaces you want to protect are taken care of. However, the decision makers are going to tell you that you need to protect everything. Thankfully Commvault allows you to protect the entire cluster all at once. That includes the configurations that allow your containers to be deployed quickly. It’s not enough to just get the data in the container. You’ve got to get the container itself too!

All of this comes in a familiar interface. One of the key negatives about the early days of container deployment was the relative immaturity of the toolset. When you’re breaking new ground you build the tool you need at the time and worry about the rest later. User interface? Who has time for that? Just hack a script together and go! Sadly for the enterprise this perspective doesn’t reassure the decision makers. On the other hand, telling them that the tried-and-true data protection platform that they’ve been using for years also supports backing up that new thing the developers are playing around with? That’s a much better way to approach the executives.

If Kubernetes represents the new frontier of cloud application deployment, Commvault represents the reliable way to keep it all safe. Your decision makers are going to feel much more comfortable making new deployment decisions on Kubernetes knowing they can rely on Commvault to back up as much or as little as they need. The reliability of the platform means assurance that it’s not going to fail because of bad infrastructure protection. That’s the kind of confidence that everyone in your enterprise wants to see.


© Gestalt IT, LLC for Gestalt IT: Protecting your Kubernetes Stack Takes Confidence with Commvault

View Details

increased Focus on Going GreenThese days you don’t have to look far to understand the importance and responsibility we all have regarding sustainability to reduce global warming and its impact on our planet. The need for change in sustainability is perhaps even more apparent at present, with world leaders meeting at the COP27 summit in Egypt, and eyes on businesses’ environmental policies across the globe. Europe is currently amid an unprecedented energy crisis with rising costs, resulting in everyone needing a better handle on energy usage. In that background, organisations must prove their green credentials to win new business with clients.

Meeting with Pure StorageI was lucky enough to meet with Prakash Darji, General Manager of the Digital Experience Group at Pure Storage, to talk about their latest sustainability initiative – the Pure1 Sustainability Assessment. Meeting with Prakash is always interesting as I know the conversation will be more than technology alone. It’ll about the coming together of technology with other business-led factors. I always enjoy these subjects where you can see technology reaching out into the business context. The intersection between business and technology is precisely where the Pure1 Sustainability Assessment comes in.

Introducing the Pure1 Storage Assessment The Pure1 Sustainability Assessment injects actionable sustainability insights directly into the Pure1 console. This information allows storage administrators and sustainability leaders to get clear insights into energy and carbon usage of the Pure platform across the organisation. With high-level expectations placed on businesses to be more sustainable, this provides critical information that can be used for reporting. The insights allow for changes to be made to improve efficiencies. The Pure1 Sustainability Assessment delivers metrics on three levels –

  1. Overall energy and carbon usage across the entire organisation
  2. Consumption at a data center location level
  3. Consumption at an individual array level

The Pure1 Sustainability Assessment can be quickly accessed from the assessment section of the dashboard menu within the Pure1 console. It initially presents a colour-coded breakdown covering all the appliances. Green indicates appliances with a good sustainability status, yellow indicates those with suggested recommendations, and red indicates appliances that require action. Quickly hovering over the appliances reveal further insights including appliance name, model and location, The Pure1 Sustainability Assessment can be quickly accessed from the assessment section of the dashboard menu within the Pure1 console. It initially presents a colour-coded breakdown covering all the appliances. Green indicates appliances with a good sustainability status, yellow indicates those with suggested recommendations, and red indicates appliances that require action. Quickly hovering over the appliances reveal further insights including appliance name, model and location, as well as the title for any insights. Clicking on the appliance shows further information detailing the insights.

In the image, an appliance has been identified with two insights – a recommendation to upgrade to newer, more efficient hardware, and an underutilisation warning. Based on their needs, users can decide if taking the recommended action would work for their technical and business needs to allow for an improved sustainability score. That is a critical element to consider when reviewing the sustainability score. Changes will need to be made based on insight from multiple parties, similar to the FinOps approach in this article.

Next to the appliance overview, is the annual power usage projection with a power-saving metric. Users also get a tuneable carbon usage meter based on their organisation’s renewable energy utilization.

Focusing on Sustainability at a Data Center LevelUsers can get a datacenter focused view by selecting the map view at the top of the page or the datacenter view icon on the right side of the page. The image below shows both views on the screen. The map view shows the geographical location and the number of appliances on a map.

From here, users can drill into the individual locations. On the right is the broader breakdown at the datacenter level. This breakdown gives users key metrics like total peak power and heat, total actual power and heat, alongside the number of rack units in use. This view offers a high-level status of each of the arrays with the ability to drill down further.

Drilling further into the information gives an individual breakdown of each appliance at the bottom of the page. From here, users can see configuration, versioning, load, power and heat metrics, and the efficiency score.

One area worth highlighting is the ability to automate report creation. Usually, those responsible for sustainability within an organisation aren’t the same as those managing the appliances. Automating the report creation allows those in need to have quick access to the required information. This allows the required business and technology discussions to be undertaken where relevant. Reports can be preconfigured and scheduled, and sent via email to the required participants.

Pure Storage has already well-documented its focus and goals regarding the sustainability of its products. For example, a recent ESG report found that Pure Storage products use less power, space, and cooling than existing and competitive solutions. I believe including a proactive tool, such as the Pure1 Sustainability Assessment, puts vital information and insight in the hands of people who can make the changes.

Final ThoughtsThis is only the beginning of this tool within the Pure ecosystem. In the future, there’ll be more options allowing users to opt into a carbon offsetting programme or offer intelligent movement based on business and performance needs and renewable energy availability. Whilst these features are not included today, it’s not hard to see such features now being a possibility at some point in the future.

Pure Storage’s FinOps, and AIOps approaches tie neatly into the subject of Sustainability Assessment. From a FinOps perspective, this tool could give further information regarding energy costs and green credentials that could be used within an organization’s FinOps assessments. AIOps could automate the movement where desired, based on energy cost or green energy availability.

The Pure1 Sustainability Assessment and the already green credentials of Pure Storage’s offerings are exactly what some organisations are looking for when making technology selections. Increasingly, organisations need to ensure that technology investments deliver tangible business outcomes, and being able to help prove or improve green credentials will be precisely the areas that will be considered.


© Gestalt IT, LLC for Gestalt IT: Improve Your Sustainability With Pure Storage

View Details

When evaluating enterprise storage solutions, organizations are often looking at flexibility as one of the decision factors. Be it in terms of price, capacity, or ability to accommodate diverse IT strategies, organizations of all sizes seek storage solutions that not only deliver immediate performance but are also capable of evolving to meet a very broad range of use cases.

Enterprise-grade capabilities and storage durability at acceptable prices are sought out by businesses of all sizes: organizations want intelligent solutions that scale, adapt, and seamlessly integrate in their broader infrastructure ecosystem, while also protecting against evolving and changing threats.

Primary Storage Needs Are ChangingSome of the most common criteria used to evaluate any primary storage system have been constant over the past decades. Customers of all sizes naturally expect enterprise storage solutions to deliver performance, reliability, efficiency, flexibility, and enterprise-grade data services at competitive prices. Although important, those expectations are now table stakes.

Today’s organizations expect more from their primary storage solutions.

They must, among others:

  • support mixed workloads and environments
  • provide an intelligent, self-optimizing architecture
  • reduce the attack surface of their environment
  • identify and combat security threats
  • meet sustainability goals

Unfortunately, legacy architectures were not designed with those modern challenges in mind. Deployment models are rigid, patching is a major pain point, with either disruptive upgrades, or long wait times to get security fixes addressed. Management tools pre-date the era of rampant online threats such as ransomware, and sustainable approaches are non-existent.

Infrastructure teams must also achieve those imperatives while facing stagnating or shrinking headcount & budgets.

To meet todays and future needs of organizations, modern enterprise storage solutions should be designed around the following principles:

  • a modern, containerized operating system that enables rapid delivery of new capabilities over time and enables the evolution of the product more rapidly
  • unified storage support to holistically address diverse workloads
  • flexible deployment models with deep ecosystem integration to support workloads across locations
  • an intelligent management framework backed by machine learning algorithms
  • a native zero trust architecture embedded at the hardware and software layers
  • sustainable deployment options with forward compatibility

These principles should help build a storage solution that is future-proof, simple to manage, adapts to new workloads and threats, minimizes efforts, and maximizes results automatically, delivering tangible business outcomes and a shorter time to value.

The PowerStore Experience: Providing Continuously Modern StorageDell Technologies’ PowerStore is built around those principles, placing it in an excellent position to deliver a superior customer experience.

PowerStore is the logical evolution of enterprise storage, built upon decades of storage experience. The solution is designed to offer an adaptable architecture that supports multiple IT strategies, with integrated intelligence to easily manage and expand, and in-place modernization to get the latest innovation, without business disruption.

The platform can be scaled up with single-drive granularity, reaching up to 4.7 PBe (effective petabytes) per appliance, or scaled out in intelligent clusters to reach over 18 PBe. Scalability is automated: all the storage services are automatically configured, providing a zero-touch experience as soon as a new drive is inserted. In scale out scenarios, intelligent clustering delivers unprecedented data and application mobility, with quick resource rebalancing capabilities.

But flexibility is not just a matter of easy capacity expansion. Modern businesses also need freedom to continuously innovate across a wide variety of application ecosystems. PowerStore’s deep integrations with VMware, Kubernetes, Ansible and more simplify DevOps orchestration and management – while unique capabilities such as Dynamic AppsON with VxRail allow PowerStore to provide scalable capacity and advanced storage services for HCI multi-cloud environments.

No matter how or where the solution evolves, PowerStore keeps workloads flexible and protected with native block, file, vVols and metro area replication. These built-in data protection capabilities support mission-critical workloads such as Oracle database servers, allowing them to remain up during planned and unplanned downtime scenarios.

The solution also harnesses the power of AI with integrated intelligence that makes changes easy. Built-in ML works in the background to continuously optimize the environment, performing intelligent data reduction, intelligent load-balancing, and a dynamic resiliency engine. This self-optimized architecture auto-tunes efficiency, performance, and availability without any manual intervention, even when making rapid changes.

PowerStore also benefits from Dell Technologies’ CloudIQ, a holistic AIOps management platform that boasts a comprehensive range of capabilities such as predictive analytics, a health scoreboard, cybersecurity alerts, and more across multiple infrastructure categories including servers, storage, networking, and cloud. Automation is another key characteristic of PowerStore, with multiple DevOps integrations and deep, two-way VMware integration.

As PowerStore technology evolves and improves over time, those benefits will be seamlessly passed over to the customer, thanks to the all-inclusive software subscription.

PowerStore and Zero Trust SecurityBesides CloudIQ security capabilities, the entire Dell Technologies ecosystem is built around a security-centric architecture that was already covered in a previous article.

PowerStore systems benefit from Dell’s focus and investments in this area: the operating system is developed in accordance with the Dell Secure Development Lifecycle (SDL), and the hardware also implements a silicon-based security and cryptographic hardware root of trust (HwROT) solution, that was described with greater detail in the security-focused blog post referenced above.

The development of a modern, containerized operating system closes the security circle by shortening software development and release times, providing better structured code, and therefore better auditability and tracking of potential issues in accordance with the Dell SDL principles.

Sustainable deployment optionsThe solution is future-proof and meets sustainability goals thanks to a flexible architecture that includes headroom for future innovations and is designed to support the next generation of storage controllers.

PowerStore brings significant improvements in terms of sustainability compared with previous generations of storage or even competitor offerings. By unifying capacity and scalability possibilities across the entire product line, customers can select CPUs based on their performance requirements, without being constrained by scalability limits.

To future proof investments in PowerStore, Dell also proposes its Anytime Upgrade program. This program provides unmatched flexibility to upgrade and scale their PowerStore investment. When a new PowerStore generation is released, the program allows customers to upgrade their existing PowerStore controller nodes into the next generation equivalent of hardware, through a simple and flexible data in-place upgrade. Customers also have the choice to upgrade to the next higher model for more performance or redeem a scale-out credit to expand their cluster with another appliance through Anytime Upgrade’s Select tier.

ConclusionEven if primary storage is a mature market, customers are still looking for innovative solutions that solve today’s challenges: regardless of their size, organizations expect versatile solutions capable of covering a broad spectrum of workloads with flexible deployment options, covering core and edge datacenters. With limited personnel and resources, customers want best-in-class solutions that are simple to deploy, operate, and maintain, while preserving investments.

Dell PowerStore proposes a comprehensive enterprise storage platform that meets and exceeds those requirements thanks to a flexible and secure architecture backed by one of the best AIOps management solutions currently available on the market. The deployment options, configurations available and rich data services & feature set position PowerStore as a polyvalent, enterprise-grade solution that will be deployed across a very large spectrum of use cases and locations, from edge locations to core data centers.


© Gestalt IT, LLC for Gestalt IT: Dell PowerStore: A Future-Proof Modern Storage Solution

View Details

Listening to Alastair Cooke’s presentation at the Intel Tech Field Day Showcase event, reminded me of a salient point when planning a cloud adoption: hardware still matters. Despite our best efforts to put multiple layers of abstraction between ourselves and the physical hardware, ultimately all software needs hardware to run on. And that hardware will impact the functionality of the software.

Flexibility in Your HardwareDepending on how you intend to consume the cloud, the degree of control you have over the hardware varies. Generally speaking, Infrastructure as a Service will provide the highest level of flexibility when it comes to hardware selection and configuration. Moving to Platform as a Service shifts the burden onto the service provider to properly vet, test, and tune the underlying hardware and systems that support the platform. Software as a Service abstracts your interaction with the hardware further still, to the degree that you may not even know what hardware and systems support the service you are consuming. For commodity off-the-shelf applications, it is desirable and economical to offload the management of hardware to a PaaS or SaaS provider. They are in the explicit business of making that application run effectively and efficiently. But in cases where you are dealing with in-house applications that are core to your business, you will find that hardware plays a vital role.

Cloud takes the existing hardware paradigm of your on-premises datacenter and flips it on its head. Consider the constraints you deal with when deploying your application on-premises. You are likely limited in the type of hardware made available and the specific combinations you can select. Even functioning as a private cloud, most enterprises have a handful of t-shirt sizes to choose from for your virtual machines. Experimenting with esoteric hardware, like a GPU card or persistent memory, is usually out of the question. The upshot is that your application has to be tuned to run on the available hardware, rather than selecting the hardware best suited to run your application.

Finding FreedomNow consider how the public cloud has removed those restrictions. No longer do you have to tune your application to run on older generation hardware or deal with a suboptimal combination of RAM and CPU. You aren’t limited to the network, storage, and compute solutions available in your enterprise datacenter. While the public cloud still has t-shirt sizes for VMs, there is an embarrassment of options. Your private cloud might have only small, medium, and large VMs. Moving to the public cloud expands you to small, medium, and large in 100 different colors and 100 different styles, each tested and maintained by the service provider. The only restriction now is that of cost, and even that is negotiable.

Having the freedom to experiment with exotic hardware and new combinations unlocks the ability to find the optimal configuration for your application. Even if you don’t plan to run the production instance of the application in the public cloud, it still serves as a proving ground for discovering the perfect hardware combination to deploy on-premises. When you need to make your capital expenditure request for new on-premises equipment, you can justify a non-standard configuration with hard data gleaned from public cloud experimentation.

ConclusionWe tend to think of the cloud as abstracting hardware away from the user, but the opposite can be true. Embracing public cloud for experimentation brings more and varied hardware closer to the consumer to help them transform their applications. While you may want to farm out non-core applications to SaaS and PaaS offerings, your core applications will benefit tremendously from being replatformed onto public cloud IaaS hardware for experimentation and possibly permanent hosting.


© Gestalt IT, LLC for Gestalt IT: Hardware Still Has to Matter

View Details

In this article presented by Solidigm, Justin Warren discusses how four corners is a great starting point, but many will more out of it by expanding on it with good architecture choices.


© Gestalt IT, LLC for Gestalt IT: Going Beyond the Four Corners of Data Storage

View Details

In this article presented by Hammerspace, Chris Evans discusses how Hammerspace is addressing the challenges of Hybrid Cloud Storage.


© Gestalt IT, LLC for Gestalt IT: True Hybrid Cloud Storage is Real and Available Today

View Details

Ned Bellavance interviews HYCU Senior Vice President of Products Subbiah Sundaram for Gestalt IT about the company's data protection solution built from the ground up with the cloud-native operational model in mind.


© Gestalt IT, LLC for Gestalt IT: Cloud Intelligent Data Protection

View Details

In this article presented by RackTop Systems, Zoë Rose discusses how RackTop BrickStor is built with security in mind focusing on user experience.


© Gestalt IT, LLC for Gestalt IT: Cyber Hygiene: Embedding Controls and Identifying Normal before Incidents Take Place?

View Details

In this article presented by RackTop Systems, Andy Banta discusses RackTop BrickStor SP's active data protection and exhaustive set of data services.


© Gestalt IT, LLC for Gestalt IT: Secure Storage with Flexible Deployment??

View Details

Pure Storage began as a provider of popular midrange all-flash storage arrays, expanded into unstructured data, and is now boldly venturing outside the traditional storage market. This is the challenge and the opportunity for the company, and one that was obvious during their Pure Accelerate techfest22 event in Los Angeles. From the FlashBlade//S to Portworx Data Services, Accelerate showed Pure looking to new markets while still focused on their traditional buyers.


© Gestalt IT, LLC for Gestalt IT: How Pure Storage Will Accelerate Into New Markets

View Details

In this article presented by Progress, Jason Gintert discusses how Flowmon uses Flowmon Collector and Flowmon Probe to secure data at the heart of a network.


© Gestalt IT, LLC for Gestalt IT: Monitoring Secure Networks with Flowmon & WhatsUp Gold

View Details

In this article presented by Progress, John Herbert discusses why Flowmon is network monitoring for a new generation.


© Gestalt IT, LLC for Gestalt IT: Progress Flowmon is a Triple Threat for Network Traffic Monitoring

View Details

In this article presented by Progress, Girard Kavelines explains why Flowmon is the triple threat for network traffic monitoring.


© Gestalt IT, LLC for Gestalt IT: Pure Gold – Progress WhatsUp Gold is Network Monitoring for A New Generation

View Details

In this article presented by Progress, Peter Welcher discusses the benefits and capabilities of WhatsUp Gold and Flowmon.


© Gestalt IT, LLC for Gestalt IT: Progress: The WhatsUp Gold and Flowmon Integration

View Details

In this Tech Field Day Showcase presented by Progress, Andy Redman and Bach Radonic Demo their Flowmon and WhatsUpGold Network Monitoring systems.


© Gestalt IT, LLC for Gestalt IT: Demos of Flowmon and WhatsUp Gold by Progress