| George Crump, of Storage Switzerland posted an article titled, Cloud Storage Reality, where he talked about the emerging class of "Cloud Storage" solutions. His conclusion: That Cloud Storage is a reality and ready for prime time. But is it? Or more specifically, is all that is called "Cloud Storage" ready for prime time. George's listing of the key advantages of cloud storage when compared with traditional enterprise storage systems, in dispersion, nodes, scale, granular, ease and self-upgrading are dead-on. | "Say, What's that mountain goat doing clear up here inthis cloud bank?" |
Similarly, we agree with his taxonomy of the three different deployment models, Service Only, Software Only and Pre-packaged Cloud. But there is a key distinction between different ways that clouds can be deployed that can make the difference between a high-risk failure and a low-risk success: Storage Just in the Cloud? While cloud storage is a proven architecture, pure Internet-based storage remains risky. Before enterprises will be willing to trust their data and their business to a provider, they first look for industry maturity, stability and reliability. After all, the pure internet-based storage industry is still in early stages of adoption, and one can argue that it already failed once, during the "Storage Utility Provider" craze at the beginning of the decade. Heck, even Enron was getting into that business. And enterprise uptime is only half the QoS battle — Even if the remote storage service provider has 100% uptime, access to the provider is limited by the reliability of the Internet networks, and access is restricted by the bandwidth to the Internet. After all, despite all the talk of bandwidth being free, the costs of an OC-3 to the Internet still makes most CFO's reach for their chests. Then, if you really want to kill pure Internet-based storage, get the lawyers involved... What Really Works Cloud Storage is production ready and widely deployed, but only in configurations that extend into the customer's data centre. I would wager that virtually all enterprise-class cloud storage deployments include data being stored in the customer's data centre. You see this with profiles of Amazon's S3 customers, and we see this with our customers. This is to be expected, of course, since all private cloud deployments exist primarily within the customer's data centre. So, to summarize, where does internet-resident cloud storage work? Cloud Storage providing off-site protection copies for data that is also held on-site. Cloud Storage providing lower-cost storage for data where high levels of QoS are not required. Cloud Storage facilitating data sharing across sites. | |
Showing posts with label S3. Show all posts
Showing posts with label S3. Show all posts
2009-02-17
Watch for Goats in the Cloud
2009-01-23
When to use S3
In response to my previous post, jeredfloyd of Permabit asked about when S3 would be useful use as storage for our customers.
These are good questions, and I'm going to elaborate on these concerns and where we see S3 as providing value to our customer.
The Bottom Line
I wouldn't use or recommend S3 for anything other than a low-grade secondary replica location for redundancy purposes. Having said that, the levels of reliability and accessibility that I've seen are already higher than what my experiences have been with tape libraries.
Bring Your Own Security
From a security standpoint, I wouldn't put anything on S3 that hasn't been encrypted and wrapped with an integrity verification layer, as we do in StorageGRID. And if the data is encrypted, there is less of a concern about deleting it if you can't get to it any more. Just throw away the keys.
As you can't implement secure wipe using their API, even if you overwrite the data, so you would also want to be sure that you're not storing really sensitive information there, even with today's standard encryption algorithms and key strengths.
S3 Isn't Inexpensive
One of the things that I want to emphasize is that based on our analysis of their economics, if you are storing data for long periods of time, it's far cheaper to just add storage nodes with SATA shelves.
Tape isn't cheaper until you're looking at 50+ TB libraries. For infrequently accessed data and redundancy copies (you need to make more when putting them on tape, since it's not as reliable as disk), it quickly becomes very economical for large capacity deployments.
Despite this, S3 Still Has Value
Having said this, even with these concerns, I see several situations where S3 support brings real value for our customers:
Based on these use-cases, it would be of most value to smaller IT shops with smaller systems. As you get into larger archives and storage systems (200+ TB), many of these situations will never come up.
Regardless of your size, having S3 as a choice as a storage tier gives administrators another tool to handle different situations, and that flexibility can be quite useful. Ultimately, it's up to them to decide if the costs (and bandwidth usage) makes sense for them.
Do you feel S3 has the reliability and availability for your customers today? I love the concept, but I've so far been scared off by horror stories of downtime. Also, what about security concerns?
These are good questions, and I'm going to elaborate on these concerns and where we see S3 as providing value to our customer.
The Bottom Line
I wouldn't use or recommend S3 for anything other than a low-grade secondary replica location for redundancy purposes. Having said that, the levels of reliability and accessibility that I've seen are already higher than what my experiences have been with tape libraries.
Bring Your Own Security
From a security standpoint, I wouldn't put anything on S3 that hasn't been encrypted and wrapped with an integrity verification layer, as we do in StorageGRID. And if the data is encrypted, there is less of a concern about deleting it if you can't get to it any more. Just throw away the keys.
As you can't implement secure wipe using their API, even if you overwrite the data, so you would also want to be sure that you're not storing really sensitive information there, even with today's standard encryption algorithms and key strengths.
S3 Isn't Inexpensive
One of the things that I want to emphasize is that based on our analysis of their economics, if you are storing data for long periods of time, it's far cheaper to just add storage nodes with SATA shelves.
Tape isn't cheaper until you're looking at 50+ TB libraries. For infrequently accessed data and redundancy copies (you need to make more when putting them on tape, since it's not as reliable as disk), it quickly becomes very economical for large capacity deployments.
Despite this, S3 Still Has Value
Having said this, even with these concerns, I see several situations where S3 support brings real value for our customers:
- If you're really small (less than 50 TB), adding storage capacity is still pretty expensive as a percentage of your yearly budget because our customers typically add in 10TB or larger increments. Using S3 as an overflow pool (keeping one or two copies locally on disk, and using S3 as your second or third copy) lets you defer that purchase for a little while, and when you do make that purchase, you can automatically migrate all the data on S3 off onto your new storage resource.
- If it takes you too long to purchase hardware, or your budgetary cycle for capital purchases takes too long, or even just an unexpected load where there just isn't time to provision more storage, you can shift second or third copies off onto S3 to free up space, and expense it to the business as a opex or project cost.
- If you have a short-term storage need, and don't want to invest in hardware yet, just put it off onto S3. It will cost a little more per TB, but since wouldn't be able to amortize the storage costs of in house hardware across the typical three-year lifespan of that hardware, it works out to be cheaper in the end.
- If you're almost full, you've ignored the alarms telling you that you don't have enough space on other nodes to repair your storage redundancy if you loose a node, and you don't have any storage ready to replace a failed storage, S3 would be a good "last resort" option for creating new replicas to restore your desired level of redundancy.
Based on these use-cases, it would be of most value to smaller IT shops with smaller systems. As you get into larger archives and storage systems (200+ TB), many of these situations will never come up.
Regardless of your size, having S3 as a choice as a storage tier gives administrators another tool to handle different situations, and that flexibility can be quite useful. Ultimately, it's up to them to decide if the costs (and bandwidth usage) makes sense for them.
Some Notes on Amazon S3
During our recent meetings, there were a fair number of questions and discussions about the economics of public cloud storage providers, such as Amazon's S3 service.
This YCombinator discussion has lots of good information about pricing, usage and experiences of some of S3's supporters and detractors. It's well worth reading.
Interestingly, thanks to a new S3 user space file-system FUSE module, Bycast has pretty much everything we need to provide a S3 tier of storage to our StorageGRID customers. Of course, such a capability would need to be productized, which would allow an administrator to have a place to configure the tier and securely enter and store their S3 credentials through our administrative interface, but thanks to the filesystem virtualizing the S3 API, all the hard work is already done.
This YCombinator discussion has lots of good information about pricing, usage and experiences of some of S3's supporters and detractors. It's well worth reading.
Interestingly, thanks to a new S3 user space file-system FUSE module, Bycast has pretty much everything we need to provide a S3 tier of storage to our StorageGRID customers. Of course, such a capability would need to be productized, which would allow an administrator to have a place to configure the tier and securely enter and store their S3 credentials through our administrative interface, but thanks to the filesystem virtualizing the S3 API, all the hard work is already done.
Cloud Storage Protocol Standardization
During one of the panel discussions at the SNIA Cloud Storage Summit, the topic of why standards for data exchange and system management would be beneficial for cloud storage. While there are many different advantages, one of the areas that I spent a few minutes talking about was the development efficiency that is realized as a result of standards.
When a protocol is standardized and adopted multiple vendors as a way to connect systems or subsystems, the following things start to emerge:
It's been my observation that the software developers and architects tend to have a major say in the selection of protocols, especially for subsystem interconnects, and that they tend to choose the protocol that makes their life the easiest. Thus, protocols that have all of these resources widely and inexpensively available quickly become the protocol of choice, resulting in a continued upward spiral of adoption, experience, tools and systems.
We're starting to see this with XAM, with #1, #4 and #5 already available, and #2, #6 in progress. And I'm sure that somewhere out there, someone's writing a book about XAM, or at least a chapter about it.
When a protocol is standardized and adopted multiple vendors as a way to connect systems or subsystems, the following things start to emerge:
- A formal protocol specification
- Web pages describing the protocol
- Books about or with chapters about the protocol
- Example open-source implementations
- Standard interface libraries
- Conformance test suites
- Benchmarking suites
- Protocol analysers and recorders
It's been my observation that the software developers and architects tend to have a major say in the selection of protocols, especially for subsystem interconnects, and that they tend to choose the protocol that makes their life the easiest. Thus, protocols that have all of these resources widely and inexpensively available quickly become the protocol of choice, resulting in a continued upward spiral of adoption, experience, tools and systems.
We're starting to see this with XAM, with #1, #4 and #5 already available, and #2, #6 in progress. And I'm sure that somewhere out there, someone's writing a book about XAM, or at least a chapter about it.
In the cloud storage protocol arena, Amazon's S3 service has such a strong lead in this area with their S3 HTTP protocol that many of these resources have already been built, despite it being a proprietary protocol. While most other cloud storage service providers have built similar HTTP protocols, with the IP ownership restrictions around Amazon's protocol still up in the air, there is a fair bit of uncertainty if their protocol will ever be able to be used with more than S3.
Which leads us back to the need for standardized protocols.
Subscribe to:
Posts (Atom)
