Dark Reading is part of the Informa Tech Division of Informa PLC

This site is operated by a business or businesses owned by Informa PLC and all copyright resides with them.Informa PLC's registered office is 5 Howick Place, London SW1P 1WG. Registered in England and Wales. Number 8860726.

Application Security //

Database Security

5/12/2008
04:45 PM
Connect Directly
Twitter
RSS
E-Mail
50%
50%

Deduplication Checklist

Here are five key questions to ask before committing to a data deduplication system.

No doubt about it: data deduplication can be a magic bullet for backup. Organizations that apply it intelligently will see faster backups, easier restores, and a reduction in power, space, and cooling costs. But put in the wrong solution, and you may instead find yourself walking the unemployment line.

Nowhere in IT does the phrase "Your mileage may vary" apply more than with data deduplication. Data reduction ratios vary depending on the type of data being backed up, the rate at which data changes between backups, and the backup scheme used.

To help companies choose the best technology for their needs, I've identified five key questions to ask:

InformationWeek Reports

Where to deduplicate?
Organizations looking to bring sanity to remote-office backups should consider remote-office/back-office backup software such as Asigra's Televaulting or EMC's Avamar that deduplicates at the source server, reducing the bandwidth needed to back up across the WAN. Larger branches, or those with less reliable WAN connections, are better served by deduplicating appliances that can replicate globally deduplicated data.

Pure Windows shops can look at Data Storage Group's ArchiveIQ, an innovative backup program that deduplicates data at the backup server. I expect EMC, CommVault, and Symantec to add backup-server deduplication over the next year or two.

How fast do you need to back up?
While vendors like to talk about speeds and feeds, the main thing is whether a given backup device is fast enough for your needs. Vendors claim their in-line deduping targets can handle 200 GB to 800 GB an hour, and their post-processing virtual tape libraries (VTLs) have data ingestion rates of up to 34 TB an hour. But the latter then may need several hours to deduplicate the data.

In addition to overall performance, make sure to look at how fast the appliance you're considering can handle a single backup stream from your biggest backup job, and how long the deduplicating post-process will take to complete.


Howard Marks

Nowhere in IT does the phrase 'Your mileage may vary' apply more than with data deduplication
Does the technology work with your backup software?
Content-aware products rely on their knowledge of the data formats that the backup applications write in. Pair a content-aware solution with a backup application it isn't equipped to manage, and you won't get any deduplication.

What interface?
Deduplicating targets come with network-attached storage and/or VTL interfaces. The NAS interface makes it easier to manage data--after all, you can't delete part of a tape, real or virtual. The problem with NAS is that it's limited to 1-Gbps Ethernet, while VTLs run over 4-Gbps Fibre Channel hardware interfaces. If you need more than backup speeds of 300 GB an hour, VTL is the way to go.

How does the technology scale?
The first data corollary to Murphy's Law states, "Data will grow to fill all available space." As a result, no matter what backup appliance or VTL you buy today, in two or three years you'll need a bigger one.

Look for devices that can expand to at least twice their size when you buy them. Gateway devices that use storage area networks and appliances that can be expanded by adding drive trays are more flexible than standalone devices. NEC's Hydrastor, using a grid architecture of accelerator and storage nodes with essentially no maximum capacity, is especially well-suited to those with fast-growing or unpredictable needs.

Illustration by John Hersey

Return to the story:
With Data Deduplication, Less Is More

Comment  | 
Print  | 
More Insights
Comments
Newest First  |  Oldest First  |  Threaded View
Aviation Faces Increasing Cybersecurity Scrutiny
Kelly Jackson Higgins, Executive Editor at Dark Reading,  8/22/2019
Microsoft Tops Phishers' Favorite Brands as Facebook Spikes
Kelly Sheridan, Staff Editor, Dark Reading,  8/22/2019
Capital One Breach: What Security Teams Can Do Now
Dr. Richard Gold, Head of Security Engineering at Digital Shadows,  8/23/2019
Register for Dark Reading Newsletters
White Papers
Video
Cartoon
Current Issue
7 Threats & Disruptive Forces Changing the Face of Cybersecurity
This Dark Reading Tech Digest gives an in-depth look at the biggest emerging threats and disruptive forces that are changing the face of cybersecurity today.
Flash Poll
The State of IT Operations and Cybersecurity Operations
The State of IT Operations and Cybersecurity Operations
Your enterprise's cyber risk may depend upon the relationship between the IT team and the security team. Heres some insight on what's working and what isn't in the data center.
Twitter Feed
Dark Reading - Bug Report
Bug Report
Enterprise Vulnerabilities
From DHS/US-CERT's National Vulnerability Database
CVE-2016-6154
PUBLISHED: 2019-08-23
The authentication applet in Watchguard Fireware 11.11 Operating System has reflected XSS (this can also cause an open redirect).
CVE-2019-5594
PUBLISHED: 2019-08-23
An Improper Neutralization of Input During Web Page Generation ("Cross-site Scripting") in Fortinet FortiNAC 8.3.0 to 8.3.6 and 8.5.0 admin webUI may allow an unauthenticated attacker to perform a reflected XSS attack via the search field in the webUI.
CVE-2019-6695
PUBLISHED: 2019-08-23
Lack of root file system integrity checking in Fortinet FortiManager VM application images of all versions below 6.2.1 may allow an attacker to implant third-party programs by recreating the image through specific methods.
CVE-2019-12400
PUBLISHED: 2019-08-23
In version 2.0.3 Apache Santuario XML Security for Java, a caching mechanism was introduced to speed up creating new XML documents using a static pool of DocumentBuilders. However, if some untrusted code can register a malicious implementation with the thread context class loader first, then this im...
CVE-2019-15092
PUBLISHED: 2019-08-23
The webtoffee "WordPress Users & WooCommerce Customers Import Export" plugin 1.3.0 for WordPress allows CSV injection in the user_url, display_name, first_name, and last_name columns in an exported CSV file created by the WF_CustomerImpExpCsv_Exporter class.