网格计算与云计算
“Cloud” Computing is 1+ yr old
Michael Sheehan’s GoGrid Blog, July 25, 2008
Confused?
SaaS
Utility Computing
SaaS = Software as a Service
?
?
Virtualization
Grid Computing
Cluster Computing
Cloud Computing
P2P
One can categorize each component
Cloud Computing
SaaS
Grid Computing
Cluster Computing
Utility Computing
Usage Model
Infrastructure
Virtualization
P2P
网格计算
What is a Grid?
Enable “coordinated resource sharing & problem solving in dynamic, multi-institutional virtual organizations.”
(Source: “The Anatomy of the Grid”)
JP Navarro <navarro@>
*
Virtual Organizations
JP Navarro <navarro@>
TeraGrid
*
What is the TeraGrid?
Technology + Support = Science
NSF已投资亿美元
自2004年10月已处于生产运行阶段,目前已用高性能网络集成了每秒750万亿次计算能力、30PB存储空间和100多个学科的数据库资源。
JP Navarro <navarro@>
*
TeraGrid’s 3-pronged strategy to further science
DEEP Science: Enabling Terascale Science
Make science more productive through an integrated set of very-high capability resources
ASTA projects
WIDE Impact: Empowering Communities
Bring TeraGrid capabilities to the broad science community
Science Gateways
OPEN Infrastructure, OPEN Partnership
Provide a coordinated, general purpose, reliable set of services and resources
Grid interoperability working group
JP Navarro <navarro@>
*
TeraGrid Used
JP Navarro <navarro@>
*
TeraGrid PI’s By Institution
TeraGrid PI’s
Blue: 10 or more PI’s
Red: 5-9 PI’s
Yellow: 2-4 PI’s
Green: 1 PI
JP Navarro <navarro@>
*
ANL/UC
IU
NCSA
ORNL
PSC
Purdue
SDSC
TACC
Computational Resources
Itanium 2 ( TF)
IA-32 ( TF)
Itanium2 ( TF)
IA-32 ( TF)
Itanium2 ( TF)
SGI SMP ( TF)
Dell Xeon ()
IBM p690 (2TF)
Condor Flock ()
IA-32 ( TF)
XT3 (10 TF)
TCS (6 TF)
Marvel SMP
( TF)
Hetero ( TF)
IA-32 (11 TF)
Opportunistic
Itanium2 ( TF)
Power4+ ( TF)
Blue Gene ( TF)
IA-32 ( TF)
Online Storage
20 TB
32 TB
1140 TB
1 TB
300 TB
26 TB
1400 TB
50 TB
Mass Storage
PB
5 PB
PB
PB
6 PB
2 PB
Net Gb/s, Hub
30 CHI
10 CHI
30 CHI
10 ATL
30 CHI
10 CHI
10 LA
10 CHI
Data Collections
# collections
Approx total size
Access methods
5 Col.
> TB
URL/DB/ GridFTP
> 30 Col.
URL/SRB/DB/ GridFTP
4 Col.
7 TB
SRB/Portal/OPeNDAP
>70 Col.
>1 PB
GFS/SRB/ DB/GridFTP
4 Col.
TB
SRB/Web Services/ URL
Instruments
Proteomics
X-ray Cryst.
SNS and HFIR Facilities
Visualization Resources
RI: Remote Interact
RB: Remote Batch
RC: RI/Collab
RI, RC, RB
IA-32, 96
GeForce
6600GT
RB
SGI Prism, 32 graphics pipes; IA-32
RI, RB
IA-32 + Quadro4 980 XGL
RB
IA-32, 48 Nodes
RB
RI, RC, RB
UltraSPARC IV, 512GB SMP, 16 gfx cards
TeraGrid Resources
100+ TF
8 distinct architectures
3 PB Online Disk
>100 data collections
JP Navarro <navarro@>
*
Science Gateways
A new initiative for the TeraGrid
Increasing investment by communities in their own cyberinfrastructure, but heterogeneous:
Resources
Users – from expert to K-12
Software stacks, policies
Science Gateways
Provide “TeraGrid Inside” capabilities
Leverage community investment
Three common forms:
Web-based Portals
Application programs running on users' machines but accessing services in TeraGrid
Coordinated access points enabling users to move seamlessly between TeraGrid and other grids.
Workflow Composer
JP Navarro <navarro@>
*
Gateways are growing in numbers
10 initial projects as part of TG proposal
>20 Gateway projects today
No limit on how many gateways can use TG resources
Prepare services and documentation so developers can work independently
Open Science Grid (OSG)
Special PRiority and Urgent Computing Environment (SPRUCE)
National Virtual Observatory (NVO)
Linked Environments for Atmospheric Discovery (LEAD)
Computational Chemistry Grid (GridChem)
Computational Science and Engineering Online (CSE-Online)
GEON(GEOsciences Network)
Network for Earthquake Engineering Simulation (NEES)
SCEC Earthworks Project
Network for Computational Nanotechnology and nanoHUB
GIScience Gateway (GISolve)
Biology and Biomedicine Science Gateway
Open Life Sciences Gateway
The Telescience Project
Grid Analysis Environment (GAE)
Neutron Science Instrument Gateway
TeraGrid Visualization Gateway, ANL
BIRN
Gridblast Bioinformatics Gateway
Earth Systems Grid
Astrophysical Data Repository (Cornell)
Many others interested
SID Grid
HASTAC
JP Navarro <navarro@>
OSG
(Open Science Grid)
*
Open Science Grid (OSG)
Origins:
National Grid (iVDGL, GriPhyN, PPDG) and LHC Software & Computing Projects
Current Compute Resources:
61 Open Science Grid sites
Connected via Inet2, NLR.... from 10 Gbps – 622 Mbps
Compute & Storage Elemets
All are Linux clusters
Most are shared
Campus grids
Local non-grid users
More than 10,000 CPUs
A lot of opportunistic usage
Total computing capacity difficult to estimate
Same with Storage
*
96 Resources across
production & integration infrastructures
20 Virtual Organizations +6 operations
Includes 25% non-physics.
~20,000 CPUs (from 30 to 4000)
~6 PB Tapes
~4 PB Shared Disk
Snapshot of Jobs on OSGs
Sustaining through OSG submissions:
3,000-4,000 simultaneous jobs .
~10K jobs/day
~50K CPUhours/day.
Peak test jobs of 15K a day.
Using production & research networks
OSG Snapshot
JP Navarro <navarro@>
NERSC
BU
UNM
SDSC
UTA
OU
FNAL
ANL
WISC
BNL
VANDERBILT
PSU
UVA
CALTECH
IOWA STATE
PURDUE
IU
BUFFALO
TTU
CORNELL
ALBANY
UMICH
INDIANA
IUPUI
STANFORD
UWM
UNL
UFL
KU
UNI
WSU
MSU
LTU
LSU
CLEMSON
MCGILL
UMISS
UIUC
UCR
UCLA
LEHIGH
NSF
ORNL
HARVARD
UIC
SMU
UCHICAGO
What is the Open Science Grid?
(+Brazil, Mexico, Tawain, UK)
JP Navarro <navarro@>
OSG应用
Genome sequence analysis
STAR: 5 TB transfer
(SRM, GridFTP)
Sloan digital sky survey
Earth System Grid:
O(100TB) online data
JP Navarro <navarro@>
Earth System Grid
JP Navarro <navarro@>
EGEE
(Enabling Grids for E-sciencE)
*
European Grid Initiative
JP Navarro <navarro@>
June 2, 2008
*
Archeology
Astronomy
Astrophysics
Civil Protection
Comp. Chemistry
Earth Sciences
Finance
Fusion
Geophysics
High Energy Physics
Life Sciences
Multimedia
Material Sciences
…
>250 sites
48 countries
>50,000 CPUs
>20 PetaBytes
>10,000 users
>150 VOs
>150,000 jobs/day
JP Navarro <navarro@>
June 2, 2008
*
Users and resources distribution
JP Navarro <navarro@>
*
EGEE workload in 2007
CPU: 114 Million hours
Data:
25PB stored
11PB transferred
? 17/05/08 $
LCG
(LHC Computing Grid)
*
Federico Calzolari
*
LHC - Large Hadronic Collider
GRID Tutorial - How to use LCG
4 experiments:
ATLAS Alice CMS LHCb
27 km long pipe
7+7 TeV
Federico Calzolari
*
Federico Calzolari
*
LCG - LHC Computing Grid
GRID Tutorial - How to use LCG
目前集成了33个国家的
140个计算中心。
2008年将执行1亿个计
算任务。
Federico Calzolari
*
*
Proxy certificate
Get your proxy certificate
temporary (usually 24h) certificate
depending on VO:
grid-proxy-init
voms-proxy-init -voms <VO>:/<VO>/Role=<role> -valid 1000:00
GRID Tutorial - How to use LCG
*
Federico Calzolari
*
Certificate
Install your certificate on the User Interface:
Log in into the UserInterface, copy there the file you exported, and create a directory where your certificate + private key will be stored:
mkdir ~/.globus
Convert PKCS12 file .p12 into the supported standard .pem
This operation will split your file in two files: the certificate () and the private key ()
openssl pkcs12 -nocerts -in <> -out ~/.globus/
openssl pkcs12 -clcerts -nokeys -in <> -out ~/.globus/
chmod 0400 ~/.globus/
chmod 0600 ~/.globus/
At end you should have something like:
[user@userinterface .globus]$ ls -al
-rw------- 1 user user 2008 Nov 13 16:50
-r-------- 1 user user 963 Nov 13 16:50
GRID Tutorial - How to use LCG
Federico Calzolari
*
*
Register to a VO
GRID Tutorial - How to use LCG
for generic user
*
*
JDL: Job Description Language
GRID Tutorial - How to use LCG
JOB overview:
JDL (job encapsulation)
main script
executable program
Creation
Submission
Status
Retrieval
*
*
JDL
Executable = "";
StdOutput = "";
StdError = "";
InputSandbox = {"",""}; # Input
OutputSandbox = {"","","out"}; # Output
VirtualOrganisation = "<VO>";
DataAccessProtocol = {"file","gsiftp","rfio","dcap"};
InputData = {"lfn:/grid/<VO>/<FILE>"};
OutputSE = "<SE>";
Requirements=Member("<SITE>",
&&
=="<QUEUE>");
GRID Tutorial - How to use LCG
*
*
Main script
#!/bin/sh
# Environment
date >> out2
hostname >> out2
# Get data
lcg-cp [-v] --vo <VO> lfn:<file> file:///
# Unpack input [: ,...]
tar -zxvf
# Compile source
g++ -o
chmod u+x
# Exec program
./ > out
# Pack output
tar -zcvf out out2
GRID Tutorial - How to use LCG
*
*
Submit a Job
Submit a JOB
edg-job-submit -o ID <JDL> # save JOBid on file ID
Selected Virtual Organisation name (from JDL): cms
Connecting to host , port 7772 # Resource Broker
Logging to host , port 9002
*********************************************************************************************
JOB SUBMIT OUTCOME
The job has been successfully submitted to the Network Server.
Use edg-job-status command to check job current status. Your job identifier (edg_jobId) is:
- :9000/tG3Xp2jT_58IUeXoY1GoZQ # JOBid
*********************************************************************************************
Control JOB status
edg-job-status <JOBid> [:9000/tG3Xp2jT_58IUeXoY1GoZQ]
*************************************************************
BOOKKEEPING INFORMATION:
Status info for the Job : :9000/tG3Xp2jT_58IUeXoY1GoZQ
Current Status: Waiting / Scheduled / Running / Done (Success/Abort)
Status Reason: Job successfully submitted to Globus
Destination: :2119/jobmanager-lcgpbs-cms
reached on: Sat Nov 17 22:38:34 2007
*************************************************************
GRID Tutorial - How to use LCG
*
*
Get the output
JOB output retrieve
edg-job-get-output <JOBid> [:9000/tG3Xp2jT_58IUeXoY1GoZQ]
Retrieving files from host: ( for :9000/tG3Xp2jT_58IUeXoY1GoZQ)
*********************************************************************************
JOB GET OUTPUT OUTCOME
Output sandbox files for the job:
- :9000/tG3Xp2jT_58IUeXoY1GoZQ
have been successfully retrieved and stored in the directory:
/tmp/jobOutput/<USER>_ tG3Xp2jT_58IUeXoY1GoZQ
*********************************************************************************
ls -al /tmp/jobOutput/calzolar_ tG3Xp2jT_58IUeXoY1GoZQ
-rw-r--r-- 1 calzolar cms 11 Nov 17 23:59 out
-rw-r--r-- 1 calzolar cms 133 Nov 17 23:59
-rw-r--r-- 1 calzolar cms 8 Nov 17 23:59
GRID Tutorial - How to use LCG
*
*
Job Requirements
JDL Requirements
everywhere
NO Requirements
at Pisa
Requirements=Member("INFN-PISA",);
on a queue 1 day at least long
Requirements=(>60*24);
on a site with at least 20 free CPU
Requirements=(>20);
on a site with at least 1 TB (unit:kb) local disk available
Requirements=anyMatch(, > 1000000000);
on a site with a given software locally installed
Requirements=Member(”VO-<VO>-TAG",);
GRID Tutorial - How to use LCG
*
*
Requirements TAGs
from SINICA
GlueHostOperatingSystemName: Scientific Linux CERN
GlueHostOperatingSystemRelease:
GlueHostOperatingSystemVersion: Beryllium
GlueSubClusterPhysicalCPUs: 0
GlueSubClusterLogicalCPUs: 0
GlueHostApplicationSoftwareRunTimeEnvironment:
LCG-2 LCG-2_1_0
LCG-2_1_1 LCG-2_2_0
LCG-2_3_0 LCG-2_3_1
LCG-2_4_0 LCG-2_5_0
LCG-2_6_0 LCG-2_7_0
GLITE-3_0_0 R-GMA
INFN-PISA SI00MeanPerCPU_1800
SF00MeanPerCPU_2000 MPICH
MPI_HOME_NOTSHARED AFS
VO-atlas-cloud-IT
[…]
GRID Tutorial - How to use LCG
*
*
Resources search
Query CPU / Storage available per VO
lcg-infosites --vo <VO> ce
#CPU Free Total Jobs Running Waiting ComputingElement
----------------------------------------------------------
165 1 1 0 1 :2119/jobmanager-pbs-cms
120 11 0 0 0 :2119/jobmanager-pbs-cms
192 110 0 0 0 :2119/jobmanager-pbs-cms
212 0 529 146 383 :2119/jobmanager-pbs-cms
227 5 312 222 90 :2119/jobmanager-lcgcondor-cms
15 15 0 0 0 :2119/jobmanager-lcgpbs-cms
80 43 0 0 0 :2119/jobmanager-pbs-cms
24 13 0 0 0 :2119/jobmanager-lcgpbs-cms
lcg-infosites --vo <VO> se
Avail Space(Kb) Used Space(Kb) Type SEs
----------------------------------------------------------
97470000
395467659 779205896
27664924 59878772
149180000
1 1
190040000 208
1000000000000 500000000000
1000000000000 500000000000
GRID Tutorial - How to use LCG
*
*
Resources search
Query available sites for my Job
edg-job-list-match <JDL>
Selected Virtual Organisation name (from JDL): cms
Connecting to host , port 7772
***************************************************************************
COMPUTING ELEMENT IDs LIST
The following CE(s) matching your job requirements have been found:
*CEId*
:2119/jobmanager-pbspro-cmsS
:2119/jobmanager-pbspro-cmsXS
:2119/jobmanager-pbs-cms
:2119/jobmanager-lcgpbs-cms
:2119/jobmanager-lcgpbs-cms
:2119/jobmanager-pbspro-cmsL
:2119/jobmanager-pbspro-cmsS
:2119/jobmanager-pbspro-cmsXS
:2119/jobmanager-lcgpbs-cms
:2119/jobmanager-lcgpbs-cms
[…]
:2119/jobmanager-lcgpbs-cms
:2119/jobmanager-lcglsf-cms4
:2119/jobmanager-lcgpbs-cms
GRID Tutorial - How to use LCG
*
*
Grid Monitoring
GRID Tutorial - How to use LCG
GOC Sinica
GridICE INFN
*
*
Grid Monitoring
GRID Tutorial - How to use LCG
AOB
云计算
Cloud Computing
*
Cloud Computing
Definition
Cloud computing is a concept of using the internet to allow people to access technology-enabled services.
It allows users to consume services without knowledge of control over the technology infrastructure that supports them.
- Wikipedia
*
Enterprise IT spending challenge
Source: IBM Corporate Strategy analysis of IDC data, Sept. 2007
Global Annual IT Spending
Estimated US$B 1996-2010
$0B
50
100
150
200
250
300
1996
1997
1998
1999
2000
2001
2002
2003
2004
2005
2006
2007
2008
2009
2010
New Server Spending
Server Mgt and Admin Costs
Power and Cooling Costs
Dream or Nightmare?
Seasonal Spikes
A Closer Look at Cloud Computing
Enterprise Cloud
Public Cloud
INNOVATIVE BUSINESS MODELS
End Users / Requestors
Government/
Academics
Industry
(Startups/ SMB/ Enterprise)
Consumers
An “Elastic” pool of high performance virtualized compute resources
Cloud applications enable the simplification of complex services
A cloud computing platform combines modular components on a service oriented architecture with flexible pricing
New combinations of services to form differentiating value propositions at lower costs in shorter time
Internet protocol based convergence of networks and devices
SIMPLIFIED SERVICES
Source: Corporate Strategy
Examples of Different Types of Services
Cloud Computing
Service Catalog
Datacenter Infrastructure
Virtual Client service
Web Application Service
Compute Service
Database service
Storage service
Content Classification
Storage backup, archive… service
Job Scheduling
Service
Collaboration Services
Google and Cloud Computing
Google与云计算
User Centric
Data stored in the “Cloud”
Data follows you & your devices
Data accessible anywhere
Data can be shared with others
music
preferences
maps
news
contacts
messages
mailing lists
photo
e-mails
calendar
phone numbers
investments
Google的三大法宝
Google File System(GFS) BigTable MapReduce
Google File System(GFS)
*
GFS Architecture
Google 48%
MSN 19%
Yahoo 33%
Files broken into chunks (typically 64 MB)
Master manages metadata
Data transfers happen directly between clients/chunkservers
Client
Client
Client
Replicas
Masters
GFS Master
GFS Master
C0
C1
C2
C5
Chunkserver 1
C0
C2
C5
Chunkserver N
C1
C3
C5
Chunkserver 2
…
Client
Client
Client
Client
Client
Client
GFS Usage @ Google
200+ clusters
Filesystem clusters of up to 5000+ machines
Pools of 10000+ clients
5+ Petabyte Filesystems
All in the presence of frequent HW failure
Google的三大法宝
Google File System(GFS) BigTable MapReduce
BigTable
Data model
(row, column, timestamp) cell contents
BigTable
Distributed multi-level sparse map
Fault-tolerance, persistent
Scalable
Thousand of servers
Terabytes of in-memory data
Petabytes of disk-based data
Self-managing
Servers can be added/removed dynamically
Servers adjust to load imbalance
Why not just use commercial DB?
Scale is too large or cost is too high for most commercial databases
Low-level storage optimizations help performance significantly
Much harder to do when running on top of a database layer
Also fun and challenging to build large-scale systems
BigTable Summary
Data model applicable to broad range of clients
Actively deployed in many of Google’s services
System provides high-performance storage system on a large scale
Self-managing
Thousands of servers
Millions of ops/second
Multiple GB/s reading/writing
Largest bigtable cell manages – 3PB of data spread over several thousand machines
Google的三大法宝
Google File System(GFS) BigTable MapReduce
MapReduce
A simple programming model that applies to many data-intensive computing problems
Hide messy details in MapReduce runtime library
Automatic parallelization
Load balancing
Network and disk transfer optimization
Handle of machine failures
Robustness
Easy to use
MapReduce Programming Model
Borrowed from functional programming
map(f, [x1,…,xm,…]) = [f(x1),…,f(xm),…]
reduce(f, x1, [x2, x3,…])
= reduce(f, f(x1, x2), [x3,…])
= …
(continue until the list is exhausted)
Users implement two functions
map (in_key, in_value) (key, value) list
reduce (key, [value1,…,valuem]) f_value
JP Navarro <navarro@>
MapReduce – A New Model and System
Two phases of data processing
Map: (in_key, in_value) {(keyj, valuej) | j = 1…k}
Reduce: (key, [value1,…valuem]) (key, f_value)
JP Navarro <navarro@>
MapReduce Version of Pseudo Code
JP Navarro <navarro@>
Example – WordCount (1/2)
Input is files with one document per record
Specify a map function that takes a key/value pair
key = document URL
Value = document contents
Output of map function is key/value pairs. In our case, output (w,”1”) once per word in the document
Example – WordCount (2/2)
MapReduce library gathers together all pairs with the same key(shuffle/sort)
The reduce function combines the values for a key. In our case, compute the sum
Output of reduce paired with key and saved
MapReduce Framework
For certain classes of problems, the MapReduce framework provides:
Automatic & efficient parallelization/distribution
I/O scheduling: Run mapper close to input data
Fault-tolerance: restart failed mapper or reducer tasks on the same or different nodes
Robustness: tolerate even massive failures:
. large-scale network maintenance: once lost 1800 out of 2000 machines
Status/monitoring
Task Granularity And Pipelining
Fine granularity tasks: many more map tasks than machines
Minimizes time for fault recovery
Can pipeline shuffling with map execution
Better dynamic load balancing
Often use 200,000 map/500 reduce tasks with 2000 machines
MapReduce: Uses at Google
Typical configuration: 200,000 mappers, 500 reducers on 2,000 nodes
Broad applicability has been a pleasant surprise
Quality experiences, log analysis, machine translation, ad-hoc data processing
Production indexing system: rewritten with MapReduce
~10 MapReductions, much simpler than old code
MapReduce Summary
MapReduce is proven to be useful abstraction
Greatly simplifies large-scale computation at Google
Fun to use: focus on problem, let library deal with messy details
A Data Playground
MapReduce + BigTable + GFS = Data playground
Substantial fraction of internet available for processing
Easy-to-use teraflops/petabytes, quick turn-around
Cool problems, great colleagues
Amazon Web Services
JP Navarro <navarro@>
Amazon Simple Storage Service
S3
JP Navarro <navarro@>
Amazon Simple Storage Service
$.15 per GB per month
storage
Object-Based Storage
1 B – 5 GB / object
Fast, Reliable, Scalable
Redundant, Dispersed
% Availability Goal
Private or Public
Per-object URLs & ACLs
BitTorrent Support
$.10 - $.18 per GB data transfer
$.01 for 1000 to 10000 requests
JP Navarro <navarro@>
Amazon S3 Concepts
Objects:
Opaque data to be stored (1 byte … 5 Gigabytes)
Authentication and access controls
Buckets:
Object container – any number of objects
100 buckets per account / buckets are “owned”
Keys:
Unique object identifier within bucket
Up to 1024 bytes long
Flat object storage model
Standards-Based Interfaces:
REST and SOAP
URL-Addressability – every object has a URL
JP Navarro <navarro@>
S3 SOAP/Query API
Service:
ListAllMyBuckets
Buckets:
CreateBucket
DeleteBucket
ListBucket
GetBucketAccessControlPolicy
SetBucketAccessControlPolicy
GetBucketLoggingStatus
SetBucketLoggingStatus
Objects:
PutObject
PutObjectInline
GetObject
GetObjectExtended
DeleteObject
GetObjectAccessControlPolicy
SetObjectAccessControlPolicy
JP Navarro <navarro@>
JP Navarro <navarro@>
JP Navarro <navarro@>
Amazon Simple Queue Service
SQS
JP Navarro <navarro@>
Amazon Simple Queue Service
$.10 per 1000 messages
Scalable Queuing
Elastic Capacity
Reliable, Simple, Secure
Inter-process messaging, data buffering, architecture component
$.10 - $.18 per GB data transfer
JP Navarro <navarro@>
Amazon SQS Concepts
Queues:
Named message container
Persistent
Messages:
Up to 256KB of data per message
Peek / Lock access model
Scalable:
Unlimited number of queues per account
Unlimited number of messages per queue
JP Navarro <navarro@>
SQS SOAP/Query API
Queues:
ListQueues
DeleteQueue
SetVisibilityTimeout
GetVisibilityTimeout
Messages:
SendMessage
ReceiveMessage
DeleteMessage
PeekMessage
Security:
AddGrant
ListGrants
RemoveGrant
JP Navarro <navarro@>
Amazon Elastic Compute Cloud
EC2
JP Navarro <navarro@>
Amazon Elastic Compute Cloud
$.10 per server hour
Virtual Compute Cloud
Elastic Capacity
GHz x86
GB RAM
160 GB Disk
250 MB/Second Network
Network Security Model
Time or Traffic-based Scaling, Load testing, Simulation and Analysis, Rendering, Software as a Service Platform, Hosting
$.10 - $.18 per GB data transfer
JP Navarro <navarro@>
Amazon EC2 Concepts
Amazon Machine Image (AMI):
Bootable root disk
Pre-defined or user-built
Catalog of user-built AMIs
OS: Fedora, Centos, Gentoo, Debian, Ubuntu, Windows Server
App Stack: LAMP, mpiBLAST, Hadoop
Instance:
Running copy of an AMI
Launch in less than 2 minutes
Start/stop programmatically
Network Security Model:
Explicit access control
Security groups
Inter-service bandwidth is free
JP Navarro <navarro@>
Root-level access
JP Navarro <navarro@>
Amazon EC2 At Work
Startups
Cruxy – Media transcoding
GigaVox Media – Podcast Management
Fortune 500 clients:
High-Impact, S hort-Term Projects
Development Host
Science / Research:
Hadoop / MapReduce
mpiBLAST
Load-Management and Load Balancing Tools:
Pound
Weogeo
Rightscale
JP Navarro <navarro@>
EC2 SOAP/Query API
Images:
RegisterImage
DescribeImages
DeregisterImage
Instances:
RunInstances
DescribeInstances
TerminateInstances
GetConsoleOutput
RebootInstances
Keypairs:
CreateKeyPair
DescribeKeyPairs
DeleteKeyPair
Image Attributes:
ModifyImageAttribute
DescribeImageAttribute
ResetImageAttribute
Security Groups:
CreateSecurityGroup
DescribeSecurityGroups
DeleteSecurityGroup
AuthorizeSecurityGroupIngress
RevokeSecurityGroupIngress
JP Navarro <navarro@>
Web-Scale Architecture
JP Navarro <navarro@>
GigaVox Economics
Implemented Amazon S3, Amazon EC2 and Amazon SQS in November 2006
Created an infinitely scalable infrastructure for less than $100 - building the same infrastructure themselves would have cost thousands of dollars
Reduced staffing requirements - far less responsibility for 24x7 operations
JP Navarro <navarro@>
分析展望
网络的迅猛发展
1986 年到2000年 计算机: × 500 网络: × 340,000
JP Navarro <navarro@>
网络发展的必然结果…
JP Navarro <navarro@>
网格计算与云计算的比较
异构资源
不同机构
虚拟组织
科学计算为主
高性能计算机
紧耦合问题
免费
标准化
科学界
同构资源
单一机构
虚拟机
数据处理为主
服务器/PC
松耦合问题
按量计费
尚无标准
商业社会
*
JP Navarro <navarro@>
云计算是广义网格的一种
“网格是构筑在互联网上的一组新兴技术,它将高速互联网、高性能计算机、大型数据库、传感器、远程设备等融为一体,为科技人员和普通老百姓提供更多的资源、功能和交互性服务。
*
Ian Foster, The Grid, 1998
JP Navarro <navarro@>
未来10年的科学
Science
网格计算
JP Navarro <navarro@>
未来10年的商业
Business
云计算
JP Navarro <navarro@>
网格书籍
JP Navarro <navarro@>
*
*
*
*
*
*
NSF TeraGrid Review
January 10, 2006
Charlie Catlett (cec@)
*
NSF TeraGrid Review
January 10, 2006
Charlie Catlett (cec@)
*
NSF TeraGrid Review
January 10, 2006
Charlie Catlett (cec@)
*
NSF TeraGrid Review
January 10, 2006
Charlie Catlett (cec@)
*
*
DEEP- To ensure that we are harnessing the enormous power of the TeraGrid resources as an integrated system we have the Advanced Support for TeraGrid Applications (ASTA) program that involves our distributed user support team (~20 people across 8 sites, coordinated by GIG). The ASTA program places 1/4 to 1/2 of a support person in a scientific application team for 4 to 12 months, working with that team to exploit TeraGrid capabilities. The ASTA program supports roughly a dozen teams at a time, with a goal of 20-25 teams supported per year.
WIDE- For over 2 decades the NSF high performance computing program, including TeraGrid, has served several thousand users very effectively. However, NSF alone funds tens of thousands of scientists, most of whom have computational requirements that do not frequently require supercomputers. The TeraGrid Science Gateways program is a set of partnerships with discipline-specific teams who are providing computational infrastructure for their science communities. The partnerships involve integrating TeraGrid as a computational and data management “service provider” embedded in the science-community infrastructure. Roughly a dozen web portal style science gateways are part of this program, with more being added continually. In addition, this program includes a peer-grid interoperation effort with Open Science Grid and a desktop application effort that is leveraging an NIH-funded biomedical project directed by Rick Stevens at the University of Chicago.
OPEN- TeraGrid began as an infrastructure involving four partner sites, grew to nine sites, and is currently organized as a set of eight resource providers and a central “grid infrastructure group” (GIG) providing central services and support as well as management and operations. This new structure allows TeraGrid to grow to dozens of resource provider sites, making it an “open” partnership. In addition, TeraGrid architecture is service-oriented, stressing open source standards such as are deployed with key software including the Globus Toolkit GT4, Condor, and other tools.
NSF TeraGrid Review
January 10, 2006
Charlie Catlett (cec@)
*
NSF TeraGrid Review
January 10, 2006
Charlie Catlett (cec@)
*
*
This Google Map mashup (using ) shows TeraGrid principal investigators for roughly 1,000 allocated projects as of June 2006.
Different icons represent different numbers of Pis per institution.
NSF TeraGrid Review
January 10, 2006
Charlie Catlett (cec@)
*
Having established a production facility with computational and storage resources and services, the TeraGrid team is now aggressively pursuing additional data management services, data collection services, and instrument access services.
RI – Remote Interactive Visualization, RB – Remote Batch Visualization, RC – Remote Interactive/Collaborative Visualization
NSF TeraGrid Review
January 10, 2006
Charlie Catlett (cec@)
*
*
TeraGrid leaders created the Science Gateways concept and initiative as a way to adapt the TeraGrid to the environments that users have adopted (. web portals), leveraging the work of these science communities in designing and building their infrastructure. This approach also leverages education efforts of these communities, and allows for an order of magnitude increase in the number of users served by TeraGrid. This increase is possible by changing the model of provision of resrouces and support, where the TeraGrid serves gateways as a “wholesale” supplier, working with the gateway providers who serve as the “retailers” providing tailored support to communities of users.
NSF TeraGrid Review
January 10, 2006
Charlie Catlett (cec@)
*
NSF TeraGrid Review
January 10, 2006
Charlie Catlett (cec@)
*
NSF TeraGrid Review
January 10, 2006
Charlie Catlett (cec@)
*
NSF TeraGrid Review
January 10, 2006
Charlie Catlett (cec@)
*
Close collaboration with EGEE/WLCG.
5 DOE Labs - BNL, Fermilab, NERSC, ORNL, SLAC. 65 Universities. 5 partner campus/regional grids.
>43,000 cores, 6 Petabyes disk cache, 10 Petabyes tape stores
15,000 CPU WallClock days/day. 1 petabyte data distributed/month.
100,000 application jobs/day. 20% cycles through resource sharing, opportunistic use.
*
*
*
NSF TeraGrid Review
January 10, 2006
Charlie Catlett (cec@)
*
*
*
*
*
*
NSF TeraGrid Review
January 10, 2006
Charlie Catlett (cec@)
*
由欧洲原子能研究中心投资,已投入1亿欧元。专为全球化分布式处理大型强子对撞机产生的每年15PB数据而设计。
*
*
*
*
*
*
*
*
*
*
*
*
*
*
*
*
*
*
*
*
*
*
*
*
Enter Title of Presentation Here
Google Confidential
*
*
*
*
*
*
*
*
*
*
*
*
*
*
*
*
*
*
*
*
*
*
*
*
*
*
*
*
网络 vs. 计算机性能
处理器速度每18个月翻一番
存储密度
每12个月翻一番
网络速度
每9个月翻一番
1986 to 2000
计算机: x 500
网络: x 340,000
2001 to 2010
计算机: x 60
网络: x 4000
光速每秒30万公里,30公里*80亿/3600=6600万公里/秒,是光速的200倍以上。
*
网络 vs. 计算机性能
处理器速度每18个月翻一番
存储密度
每12个月翻一番
网络速度
每9个月翻一番
1986 to 2000
计算机: x 500
网络: x 340,000
2001 to 2010
计算机: x 60
网络: x 4000
光速每秒30万公里,30公里*80亿/3600=6600万公里/秒,是光速的200倍以上。
*
*
*
*
*
*
*