You are viewing an old version of this page. View the current version.

Compare with Current View Page History

« Previous Version 8 Next »

Introduction

Purpose

This is a draft spec for utilizing Q-in-Q to provide scalable isolated networks in the cloudstack environment.

Q-in-Q, also referred to as "double-tagged" vlans, is the concept of nesting tagged vlans inside each other. Thus, each
of the 4096 available vlans can also host 4096 vlans. This provides a way to scale isolated networks that is more compatible
with standard network equipment and lower overhead than GRE or other technologies that create point-to-point tunnels.

This has been designed and tested on the KVM platform, there's no reason to believe it wouldn't work on others, but more
time and expertise will be required to create and verify code.

References

  • relevant links

Document History

Glossary

Feature Specifications

  • Quality risks (test guidelines)
    • There is some question on how to handle MTU. See the chart included for the current solution.
    • No single system will scale to thousands of networks, regardless of technology. Therefore, we need to ensure that  CloudStack
      is properly tearing down networks that it no longer needs, and creating only networks necessary for the running instances
    • The current implementation relies on naming conventions to differentiate between a tagged interface that CloudStack should
      treat normally and an interface that CloudStack should treat as a physical interface. This works on KVM since linux provides
      two standard interface names for tagged vlan interfaces, and one is rarely used. However, there could be issue in the event
      that a customer is using the rare convention. We'd need to document this in the standard CloudStack network setup for KVM
      hosts.
  • Supportability characteristics:
    • The implementation leverages CloudStack's existing network/bridging management code, so troubleshooting would be
      largely the same. The only caveat is in regards to the MTU as mentioned and covered later on.
  • Configuration

A traditional CloudStack advanced network environment might look like the below chart, with management, storage, and a public network created by the admin, and multiple vlans provisioned by cloudstack as needed:

q-in-q example advnetworking

This functional spec extends this design by allowing tagged interfaces to be utilized at the physical interface level:

q-in-q example advnet2

The admin simply needs to create any 'vlan#' devices, and CloudStack uses them as physical devices.

CloudStack doesn't support defining actual physical devices to be used; instead it uses "traffic labels" to allow the admin to specify which bridges to use for what traffic, and then cloudstack determines the physical devices from those bridges. It then uses those physical devices to create subsequent tagged interfaces/bridges dynamically. Normally, if CloudStack finds that a bridge is on a tagged interface, it then looks up the parent of that interface. A small patch simply keeps CloudStack from looking up the parent if the device is a vlan# device, so that the vlan# dev is instead treated as a physical device and subsequent tagged networks are created there.

MTU

Actual MTU values and what is included in them vary slightly between manufacturers and operating systems, but in general it should be kept in mind that Q-in-Q requires extra space in order to contain the new vlan tag. The table below provides an idea of what MTU values have been tested to work. Generally speaking, an admin should increase the MTU on the applicable switch hardware in order to be compatible with existing MTUs in CloudStack system VMs and guest instances.

Switch MTU

1500

1532

9000

9032

Instance MTU 1468*

Y

Y

Y

Y

Instance MTU 1500

N

Y

Y

Y

Instance MTU 8968*

N

N

Y

Y

Instance MTU 9000*

N

N

N

Y

* Special consideration is required for MTU != 1500 for any virtual routers. Will need to add the ability to set MTU in the VR, perhaps in the same way that SSVM gets MTU set in the global settings.

  • deployment requirements (fresh install vs. upgrade) if any
  • system requirements: memory, CPU, disk space, etc
  • interoperability and compatibility requirements:
    • OS
    • xenserver, hypervisros
    • storage, networks, other
  • list localization and internationalization specifications 
  • explain the impact and possible upgrade/migration solution introduced by the feature 
  • explain performance & scalability implications when feature is used from small scale to large scale
  • explain security specifications
    • list your evaluation of possible security attacks against the feature and the answers in your design* *
  • explain marketing specifications
  • explain levels or types of users communities of this feature (e.g. admin, user, etc)

Use cases

put the relevant use case/stories to explain how the feature is going to be used/work

Architecture and Design description

  • discussion of alternatives amongst design ideas, their resources/time tradeoffs and limitations. Explain why a certain design idea is chosen over others
  • highlight architectural patterns being used (queues, async/sync, state machines, etc)
  • talk about main algorithms used
  • explain what components are being changed and what the dependent components are
  • regarding database: talk about tables being added/modified
  • performance implications: what are the improvements or risks introduced to capacity, response time, resources usage and other relevant KPIs
  • preferably show class diagrams, sequence diagrams and state diagrams
  • if possible, publish signatures of all methods classes and interfaces implement, and the explain the object information of different classes

Appendix

Appendix A:

Appendix B:

  • No labels