Pages

Showing posts with label storage. Show all posts
Showing posts with label storage. Show all posts

Saturday, March 13, 2021

Storage Basics - Object, File and Block

In Simple terms, Object storage is a computer data storage mechanism that manages data as units called objects, as opposed to other storage mechanisms like file storage where data is stored as file hierarchy and block storage which manages data as blockers within sectors and tracks. In this article, we will discuss more details about the storage mechanisms and see how they work.


Object Storage

Object Storage is a collection of data with one unique identifier and amount of metadata that is stored as Objects. Data is managed as units called blocks.


Data : The data can be anything that makes up the object. It can be anything from audio file to photo album.


Identifier : Data that is added to the object storage will get a UUID ( Universally unique identifier )and GUID ( Globally unique identifier ). These 2 identities are unique and 128bit long. These identifiers are very useful when accessing the data from a large set of object storage.


Metadata : this is data about the data or we can call them as labels to data. Metadata is data attached to the original data. This data can be any information that is used to classify or identify the data. Metadata can be taught as labels for the data.


Advantages & disadvantages : The primary advantage using this object storage is the huge amount of data that can be stored. Since the data is unstructured, there is lots of data that can be stored but yet provides an easy way of accessing the data. So though we have large unstructured data stored, we can still access that quite easily. Huge amount of data storage is achieved due to its flat structure - by using GUIDs instead of hierarchies structure like file and block.


The Data is easily accessible with also using the metadata that we attach to the data. This metadata is quite customizable and expanded thus allowing more easier access to the data being stored in object storage.


Backup and Archiving : Since data is quite unstructured, performing backup and archiving the data is quite easy and fast. 


Advantages for object storage include:

Greater data analytics : Since data is driven by metadata and deep level of classification attached to that, analysis or accessing of data is Very good

Infinite scalability : Add as much as data and there is no limit

Faster data retrieval :Since lot of label and metadata attached to data and no specific structuring way , data can be accessed very fast 

Reduction in cost : Cost is very low for storing data

Optimization of resources : Since lot of data is being stored with no limits, resources can be efficiently utilized


Data that benefits the most from object storage includes:

Unstructured data such as music, images, and videos

Backup and log files

Large sets of historical data

Archived files

 

Tools available : Amazon S3 bucket, Microsoft Azure Blob


File Storage

File storage is the simplest storage available to us. In file storage, the data is stored in files. These files in turn are organized in folders or directories in a hierarchical fashion. To access a file, users or machines only need the path from directory to subdirectory to folder to file.


Advantages & disadvantages : Since data is being stored in files and folders, it's quite easy to organise and access the data. The major advantage of using this file storage is sharing and security. The file storage can be shared with multiple people and anyone can set permissions on the files and folders in file storage. This helps in sharing, securing, collaborating and accessing files and folders.


The primary disadvantage with this file storage is, however we plan to increase the data in files and folders at some point it will be very complex to handle the sharing, permissions and security. Things will get more complex with more and more file storage. 


In contrast to block storage, a system with file storage does not take the data of the file apart. The file is stored as a whole and called up again in this form. File level storage other than built in harddrives, we have 2 other types 

Network attached Storage ( NAS ) : Storage system connected to a network and available to all participants of the network. 


Direct Attached Storage ( DAS ) : Storage system directly connected to the Computer in the form of a External System

The other major advantage is the inexpensive storage drives. If you want more storage we can attach the external drive or network drives as file systems to the current machine. The pricing of these drives is very low when compared. 


The major disadvantage is the rising complexity with growing file systems. When the file system grows, managing it will be complex. File storage can be when we need

Local file sharing

Centralized file collaboration

archiving/storing

Backup/disaster recovery


Tools Available : Amazon Elastic File System, Azure Files


Block Storage

The final storage type is the block storage and currently favorite for many cloud based applications. In this type, data is broken into pieces called blocks and then stored across a system that can be physically distributed to maximize efficiency. Each block will have a unique identifier which allows the storage system to put these blocks together when data is needed. Data is stored in fixed sized blocks and a unique address serves as a metadata identifying each block


Advantages & disadvantages

The major advantage with this type of storage is the ability to quickly retrieve and modify data that is spanned across locations. Block storage divides the data into blocks and span them across different environments thus creating multiple paths to data. This helps in retrieving the data faster when required. When a user requests for data, the underlying Operating system gathers the blocks and reassembles them into one data block and provides that to the application. The Server operating system will be responsible for separating storage by fixed sized blocks, spanning them to different environments, reassembling them when needed. This reassembling will be done by using the Server address of blocks. Protocols like Fiber Channel over ethernet(FCoE) , Internet Small Computer System Interface ( iSCSi) etc used to access the block storage data. These are commonly used in a storage area network (SAN) where high performance is required with High I/O and low latency


The primary disadvantages are the ability to add more metadata to the blocks. The other challenge is that this block storage cannot be accessed by multiple participants at the same time unlike File storage.


Tools Available : Azure Managed Disk, Aws Elastic Block Storage


Hope this helps you to understand the basics storage mechanisms

Read More

Sunday, October 7, 2018

Distributed File System - GlusterFS


There are many cases where applications would require accessing data or would require a place to upload files. To support these, we have a shared file system. Services like Nfs or Samba provides shared drives access to application running anywhere. What if we need to have high availability for the shared drives?

In the technology world, it is always crucial to keep data highly available to ensure it is accessible to every application/user. High availability of data is achieved by distributing the data across multiple nodes or multiple volumes in multiple nodes.
Client machines/users can access the storage as like local storage when mounted. The advantage is the volumes are configured as distributed. 

Imagine if the users are doing a heavy read/write operation on the same NFS volume. The Memory or cpu on the machine hosting the volume can become slow due to load. What if we can combine the memory and processing power of 2 machines and their individual discs to form a single volume accessed by the clients?. This is where the distributed file systems come into picture.

In computing, a distributed file system (DFS) or network file system is any file system that allows access to files from multiple hosts sharing via a computer network. This makes it possible for multiple users on multiple machines to share files and storage resources. So we can create 2 machines and have directories created which will be shared as single volume to the external world.

What is GlusterFS?
GlusterFs does the same thing of combining multiple storage servers to form a large, distributed drive. GlusterFs is a open source, scalable network file system suitable for high data intensive workloads such as media streaming, storage, content delivery etc
In this article we will see how we can configure GlusterFS on a Centos 7 Machines.

For this we will use 3 machines of which 2 are servers providing the volumes and other one acts as a client.
1.  Configure 3 machines with Centos 7. 

2. Add the details of the 3 machines to /etc/hosts file in all 3 machines. Below is my configuration
10.131.224.54     server1.example.com     server1
10.131.224.149   server2.example.com     server2
10.131.225.130   client.example.com        client

3. Add the extra repo details to the Centos 7 using the below commands,
wget http://dl.fedoraproject.org/pub/epel/epel-release-latest-7.noarch.rpm
rpm -ivh epel-release-latest-7.noarch.rpm

4. Create a Glusterfs repo. Create a file glusterfs.repo in /etc/yum.repos.d/glusterfs.repo with the below content

[root@manja17-I18062 ~]# cat /etc/yum.repos.d/gluster.repo
[gluster41]
name=Gluster 4.1
baseurl=http://mirror.centos.org/centos/7/storage/x86_64/gluster-4.1/
gpgcheck=0
enabled=1

5. Install and start the glusterfs, yum install glusterfs-server samba -y
    Start the service using ,  systemctl enable glusterd.service and
    systemctl start glusterd.service

6. Check the glusterfs version
[root@manja17-I18063 ~]# glusterfsd --version
glusterfs 4.1.5
Repository revision: git://git.gluster.org/glusterfs.git
Copyright (c) 2006-2016 Red Hat, Inc.
GlusterFS comes with ABSOLUTELY NO WARRANTY.
It is licensed to you under your choice of the GNU Lesser
General Public License, version 3 or any later version (LGPLv3
or later), or the GNU General Public License, version 2 (GPLv2),
in all cases as published by the Free Software Foundation.

7. Disable the firewall using, systemctl stop firewalld

8. If you use a firewall, we need to make sure Tcp ports 111, 24007, 24008,24009 are open on the server1 and server2 if we have not disabled the firewall with the above step

9. Next we need to configure the trusted pool storage. We will be adding server2 as a trusted pool to server1. For this we will run the glusterfs command from server1.

[root@manja17-I18062 ~]# gluster peer probe server2.example.com
peer probe: success.

10. Next check the status using,
[root@manja17-I18062 ~]# gluster peer status
Number of Peers: 1

Hostname: server2.example.com
Uuid: 9964b028-585d-4cf1-b67b-4bbf9be2a976
State: Peer in Cluster (Connected)

11. Now Let's create a share named testervol with  two replicas. We need to understand that the number of replicas are equal to the number of servers that we configured since we need to set up mirroring.   We will be configuring this directory on both machines with location /data and to the external world it will shown as testervol.

[root@manja17-I18062 ~]# gluster volume create testervol replica 2 transport tcp server1.example.com:/data server2.example.com:/data force
volume create: testervol: success: please start the volume to access data

Once this is done,we will see /data location created in both machines.

12. Start the Volume
[root@manja17-I18063 data]# gluster volume start testervol
volume start: testervol: success

13 . Check the volume info using
[root@manja17-I18063 data]# gluster volume info
Volume Name: testervol
Type: Replicate
Volume ID: 7196569d-2a27-4cc9-9918-acb952b42549
Status: Started
Snapshot Count: 0
Number of Bricks: 1 x 2 = 2
Transport-type: tcp
Bricks:
Brick1: server1.example.com:/data
Brick2: server2.example.com:/data
Options Reconfigured:
transport.address-family: inet
nfs.disable: on
performance.client-io-threads: off

By default all clients can access the volume defined. We need to define some access controls on who will access these.

Setting the GlusterFs Client ( on the client Machine )
1. Create a directory in the /mnt location with glusterfs
[root@manja17-I18064 ~]# mkdir /mnt/glusterfs

2. Mount the volume
[root@manja17-I18064 ~]# mount.glusterfs server1.example.com:/testervol /mnt/glusterfs

3. Check for the volume
[root@manja17-I18064 ~]# mount | grep testervol
server1.example.com:/testervol on /mnt/glusterfs type fuse.glusterfs (rw,relatime,user_id=0,group_id=0,default_permissions,allow_other,max_read=131072)

[root@manja17-I18064 ~]# df -h
Filesystem                          Size     Used    Avail   Use% Mounted on
/dev/mapper/centos-root     50G     4.1G   46G      9%     /
devtmpfs                            3.9G    0        3.9G     0%    /dev
/dev/sdb1                          60G      33M    60G      1%    /loddisk2
/dev/sda1                          1014M  161M  854M    16%   /boot
server1.example.com:/testervol   50G  4.7G   46G  10% /mnt/glusterfs

Test the volume,
On client run , [root@manja17-I18064 ~]# touch /mnt/glusterfs/test1

Check on the server1 or server2
[root@manja17-I18062 data]# pwd
/data

[root@manja17-I18062 data]# ls
test1

More to Come,Happy learning :-)
Read More