20130115

Hot-remove, Hot-add drives under Linux

UPDATE: See http://burning-midnight.blogspot.com/2013/01/hot-addremove-strangeness.html for some strangeness I encountered while doing the following...

I keep looking for this because I keep forgetting it.  Now I have two scripts that make my job a lot easier.  I also recently started using device aliasing under ZFSonLinux, meaning I can type things like "a1" and "b3" instead of scsi-1ATA_WDC_WD10JPVT-00A1YT0_WD-WXB1EA......

BUT the device aliasing has a downside; I'd still have to dig through the dev tree and match up zpool device names to their semi-real system counterparts...up till now!

Here's a script to scan every SCSI bus on the system, so that when you add a drive it should just find it (my system has 8 somehow, by the way):


for X in /sys/class/scsi_host/host?; do
  echo "- - -" > ${X}/scan
done

And here's a script I found and made one minor change to (had to fix what was maybe a typo or a difference in shells).  If you supply the exact device path (such as /dev/zpool/a5), it will hot-remove it for you.  Someone commented that calling the device by name (sda, sdb) works too, but this does not seem to be applicable where the ZFS device aliasing is concerned.  Anyway...

#!/bin/bash
# (c) 2009 by Dennis Birkholz (firstname DOT lastname [at] nexxes.net)
#
# This program is free software: you can redistribute it and/or modify
# it under the terms of the GNU General Public License as published by
# the Free Software Foundation, either version 2 of the License, or
# (at your option) any later version.
#
# This program is distributed in the hope that it will be useful,
# but WITHOUT ANY WARRANTY; without even the implied warranty of
# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE.  See the
# GNU General Public License for more details.
#
# You can received a copy of the GNU General Public License at
# .
function usage {
echo "Usage $0 [device]"
echo
echo "Disable supplied SCSI device"
exit
}
# Need a parameter
[ "$1" == "" ] &&
usage
# Verify parameter exists
( [ ! -e "$1" ] || [ ! -b "$1" ] ) &&
echo "Supplied devices does not exist or is not a block device." >/dev/stderr &&
exit 1
# Verify SCSI disk entries exist in /sys
[ ! -d "/sys/class/scsi_disk/" ] &&
echo "Could not find SCSI disk entries in sys, aborting." >/dev/stderr &&
exit 2
# Get major/minor device string of device
major=$(stat --dereference --format='%t' "$1")
major=$(printf '%d\n' "0x${major}")
minor=$(stat --dereference --format='%T' "$1")
minor=$(printf '%d\n' "0x${minor}")
deviceID="${major}:${minor}"
echo "Major/Minor number for device '$1' is '${deviceID}'..."
for device in /sys/class/scsi_disk/*; do
[ "$(< ${device}/device/block/*/dev)" != "${deviceID}" ] && continue
scsiID=$(basename "${device}")
echo "Found SCSI ID '${scsiID}' for device '${1}'..."
echo 1 > ${device}/device/delete
echo "SCSI device removed."
exit 0
done
echo "Could not identify device as SCSI device, aborting." >/dev/stderr
exit 4
I will say I'm a little disappointed that after all this time someone hasn't come up with a more well-published way to do this on systems.  Of course, most of us don't hot-swap our drives, so maybe I shouldn't be TOO disappointed.

20130110

Fixing Balance-ALB (Mode 6) Bonding for KVM

I ended up contacting the netdev list, looking to see if the problems I was experiencing with Balance-ALB were fixable and if a fix would be accepted.

Good news!  It was already fixed!

Bad news...it's only fixed in the 3.8 release candidate right now.

The responder pointed me to the patch submission that fixed the issue at hand: balance-ALB would no longer stomp MACs that did not originate from the host itself.  Simple enough to apply to the 3.0 kernel, but there had been some other changes that caused both a hunk to fail and the build to fail.  I had to pull in a function from upstream and backport it into one of the headers.  The next challenge was getting the .deb packages built...I made the mistake of doing this on a ramdrive, not realizing it would compile everything three times and generate three images.  24G of ramdrive later, it was done.

The installation, at least, was easy enough...thanks to the .debs.  After rebooting, the bond worked correctly, and the MACs for all my virtuals are now visible and correct!

For posterity, this is the link I was given for the original patch:

http://git.kernel.org/?p=linux/kernel/git/stable/linux-stable.git;a=patch;h=567b871e503316b0927e54a3d7c86d50b722d955

Below is the patch for the 3.0 kernel.  The patch appears to build for kernels up to (but not including) the 3.7 series.  3.7 should work if you omit the etherdevice.h portion of the patch.

diff -uNr linux-3.0.0-a/drivers/net/bonding/bond_alb.c linux-3.0.0-b/drivers/net/bonding/bond_alb.c
--- linux-3.0.0-a/drivers/net/bonding/bond_alb.c        2013-01-10 12:47:53.000000000 -0500
+++ linux-3.0.0-b/drivers/net/bonding/bond_alb.c        2013-01-10 12:50:58.000000000 -0500
@@ -666,6 +666,12 @@
        struct arp_pkt *arp = arp_pkt(skb);
        struct slave *tx_slave = NULL;

+       /* Don't modify or load balance ARPs that do not originate locally
+        * (e.g.,arrive via a bridge).
+        */
+       if (!bond_slave_has_mac(bond, arp->mac_src))
+               return NULL;
+
        if (arp->op_code == htons(ARPOP_REPLY)) {
                /* the arp must be sent on the selected
                * rx channel
diff -uNr linux-3.0.0-a/drivers/net/bonding/bonding.h linux-3.0.0-b/drivers/net/bonding/bonding.h
--- linux-3.0.0-a/drivers/net/bonding/bonding.h 2011-07-21 22:17:23.000000000 -0400
+++ linux-3.0.0-b/drivers/net/bonding/bonding.h 2013-01-10 12:51:05.000000000 -0500
@@ -18,6 +18,7 @@
 #include
 #include
 #include
+#include
 #include
 #include
 #include
@@ -431,6 +432,18 @@
 }
 #endif

+static inline struct slave *bond_slave_has_mac(struct bonding *bond,
+                                              const u8 *mac)
+{
+       int i = 0;
+       struct slave *tmp;
+
+       bond_for_each_slave(bond, tmp, i)
+               if (ether_addr_equal_64bits(mac, tmp->dev->dev_addr))
+                       return tmp;
+
+       return NULL;
+}

 /* exported from bond_main.c */
 extern int bond_net_id;
diff -uNr linux-3.0.0-a/include/linux/etherdevice.h linux-3.0.0-b/include/linux/etherdevice.h
--- linux-3.0.0-a/include/linux/etherdevice.h   2011-07-21 22:17:23.000000000 -0400
+++ linux-3.0.0-b/include/linux/etherdevice.h   2013-01-10 12:51:16.000000000 -0500
@@ -275,4 +275,37 @@
 #endif
 }

+/**
+ * ether_addr_equal_64bits - Compare two Ethernet addresses
+ * @addr1: Pointer to an array of 8 bytes
+ * @addr2: Pointer to an other array of 8 bytes
+ *
+ * Compare two Ethernet addresses, returns true if equal, false otherwise.
+ *
+ * The function doesn't need any conditional branches and possibly uses
+ * word memory accesses on CPU allowing cheap unaligned memory reads.
+ * arrays = { byte1, byte2, byte3, byte4, byte5, byte6, pad1, pad2 }
+ *
+ * Please note that alignment of addr1 & addr2 are only guaranteed to be 16 bits.
+ */
+
+static inline bool ether_addr_equal_64bits(const u8 addr1[6+2],
+                                           const u8 addr2[6+2])
+{
+#ifdef CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESS
+        unsigned long fold = ((*(unsigned long *)addr1) ^
+                              (*(unsigned long *)addr2));
+
+        if (sizeof(fold) == 8)
+                return zap_last_2bytes(fold) == 0;
+
+        fold |= zap_last_2bytes((*(unsigned long *)(addr1 + 4)) ^
+                                (*(unsigned long *)(addr2 + 4)));
+        return fold == 0;
+#else
+        return ether_addr_equal(addr1, addr2);
+#endif
+}
+
+
 #endif /* _LINUX_ETHERDEVICE_H */


20130107

Balance-ALB Woes...

I've concluded my research and here is the answer.

It's dead, Jim.

If your virtual's MACs are getting squashed, look no further than Balance-ALB (mode 6).

I'm not sure if this can be fixed or not, but right now it sucks.  After much testing and lots more reading, it seems this is a "known problem" and doesn't look like it's going to be fixed.  For reference, here's the configuration:

    vnet0 -> br0 -> bond0 -> eth1, eth2, ...  (note to self: make this a pretty picture)

Where bond0 is mode-6 over the listed interfaces.  ALB is supposed to balance transmits AND receives, so to accomplish this it apparently snags ARPs from the wire and replaces them with one of the several MACs of its slaves.  I think, if I recall correctly, it picks a slave in a round-robin fashion.  Anyway, the problem seems to be that when ARPs for the virtuals under the bridge come in, ALB snags those as well and scarfs them, stomps them, and sends out its own MAC.

Thus, the spice does not flow.

Symptoms include: intermittent ping, intermittent connectivity, ARP table reading one of the bond's MACs instead of the virtual's, headache, nausea, and some minor vomiting.

Nothing else, save mode-0 (which really doesn't count) does any sort of receive-side load-balancing.  What I would REALLY like to see is ALB more intelligently handle ARP requests, such that it doesn't squash those of the virtuals that are properly serviced elsewhere.  Ideally, it should not squash any ARP replies that do not have anything to do with its own physical adapters.  That seems like it ought to be a relatively easy fix...except I don't know where in the code to fix it...yet.  I might try someday...

Use Mode-4 you say?  Nope, it doesn't load-balance under the bonding driver.  Check out libteam instead and then go crazy trying to build it under Ubuntu 11.10.  It looks like it's under 12.04, except my stupid cluster doesn't run 12.04 because of issues with OCFS2 and the DLM and a lot of other bullshit that is taking way too much energy to solve.  Plus, Mode-4 is great if you're using a single switch.  I'm using two, for redundancy.  I could go back to Sins of the Bond and team two mode-4's under maybe a mode-1 (XOR), but then I'm kinda back where I started with no really awesome receive-load-balancing.

The ALB problem is especially nefarious because occasionally the virtual's REAL MAC will appear in the ARP cache.  Now, just What The Fuck is up with that?  It makes me think this issue with ALB is more of a bug than a feature.

Temporary fixes in the meantime: switch to TLB (mode 5), or any other mode that doesn't involve borking ARPs.  Strangely, even though TLB manages load-balancing on transmit, it doesn't display the same ARP-hell that ALB suffers.


20130103

Bridging, Bonding, and VM MACs...OH MY!!

I've been pouring through various swaths of documentation and forum posts all day, to little avail.  Perhaps it's buried in the "how Linux bridging works" doc somewhere in the kernel files, but I'm gonna ask here for kicks:  I've noticed a strange trend, and ordinarily I wouldn't care except I lose connectivity for short periods with my virtuals.  I have several hosts that all network by bridging virtual ethernet adapters with a bond of the physical adapters.  For sake of clarity: phy -> bond -> bridge, with mode-6.  So, I ping away at a host, and some of the pings come back "TTL Exceeded".  Not fun, especially since I also have trouble ssh'ing into the vm.  I can get to the console with VNC, and play around with it and even ping from the vm to, say, the firewall without any issues.  After some reading, the notion of MAC conflicts and bridging and bonding issues came up.

I pulled up Wireshark and started examining the ARP responses for said given host.  I detected that the ARP given for the vm's IP varied among the bonded adapters - I half expected this, since mode-6 is supposed to load-balance automatically.  However, when I started pinging FROM my vm TO my Wireshark machine, the MAC in my cache suddenly changes to the vm's actual MAC.  Once the pinging is stopped, the MAC eventually reverts to one of the physical adapters.

To sum up: pinging WS -> VM = bond MAC, whereas pinging VM -> WS = VM MAC.  It's like the VM's pings become unintentional gratuitous ARPs.

So, what's the deal here?  Is this "working right" and my problem with intermittent connectivity somewhere else?  Is it normal for either the bridge or the bond to reply with their own selective MAC address to ARP requests?  But then why bother doing that when you could just as easily let the VM publish its own response?  Obviously the MAC isn't getting nuked when the VM transmits, since it appears in my ARP cache during that ping experiment.....   ??  I get the feeling there is a setting that needs to be modified.

20121124

Broken cman init script?!

Nothing is going right in Ubuntu 12.04 for cluster-aware file systems.

OCFS2 seems borked beyond belief.
(Update 2013-02-25: I believe I have made progress on the OCFS2 front:  http://burning-midnight.blogspot.com/2013/02/quick-notes-on-ocfs2-cman-pacemaker.html)

GFS2 has its own issues, some of which I will detail here.

You CAN get these two file systems running on 12.04.  Whether or not the cluster will remain stable when you have to put a node on standby is another question entirely, and a very good question.  It shouldn't even BE a question, but it is, and the answer is a resounding FUCKING NO!  Well, at least, as far as OCFS2 is concerned.  The problems there lie in who manages the ocfs2_controld daemon.  CMAN ought to do it, but CMAN doesn't want to.  Starting it in Pacemaker causes horrible heartburn when you put a node into standby, and things just all fall apart from there.

I decided to try out GFS2.  After installing all the necessary packages, and manually running bits here and there to see things work, I could not get Pacemaker to mount the GFS2 volume.  First problem was the CLVM: if you want to be able to shut down a node without shooting the fucker, you'll need to make sure the LVM system can deactivate volume groups.  The standard method of vgchange -an MyVG doesn't work for the cluster-aware LVM.  It complains loudly about "activation/monitoring=0" being an unacceptable condition for vgchange.  This detailed in this bug: https://bugs.launchpad.net/ubuntu/+source/lvm2/+bug/833368

The solution suggested there, at least where the OCF script is concerned, works: change the lines that use "vgchange" to include "--monitor y" on the command-line, and it will magically work again.

My cluster starts a DRBD resource, promotes it (dual-primary), then starts up clvmd (ocf:lvm2:clvmd), activates the appropriate LVM volumes (ocf:heartbeat:LVM), then mounts the GFS2 file system (ocf:heartbeat:FileSystem).  These are all cloned resources.
primitive p_clvmd ocf:lvm2:clvmd \
    op start interval="0" timeout="100" \
    op stop interval="0" timeout="100" \
    op monitor interval="60" timout="120"
primitive p_drbd_data ocf:linbit:drbd \
    params drbd_resource="data" \
    op start interval="0" timeout="240" \
    op promote interval="0" timeout="90" \
    op demote interval="0" timeout="90" \
    op notify interval="0" timeout="90" \
    op stop interval="0" timeout="100" \
    op monitor interval="15s" role="Master" timeout="20s" \
    op monitor interval="20s" role="Slave" timeout="20s"
primitive p_fs_vm ocf:heartbeat:Filesystem \
    params device="/dev/cdata/vm" directory="/opt/vm" fstype="gfs2"
primitive p_lvm_cdata ocf:heartbeat:LVM \
    params volgrpname="cdata"
ms ms_drbd_data p_drbd_data \
    meta master-max="2" clone-max="2" interleave="true" notify="true" clone cl_clvmd p_clvmd \
    meta clone-max="2" interleave="true" notify="true" globally-unique="false" target-role="Started"
clone cl_fs_vm p_fs_vm \
    meta clone-max="2" interleave="true" notify="false" globally-unique="false" target-role="Started"
clone cl_lvm_cdata p_lvm_cdata \
    meta clone-max="2" interleave="true" notify="true" globally-unique="false" target-role="Started"
colocation colo_lvm_clvm inf: cl_fs_vm cl_lvm_cdata cl_clvmd ms_drbd_data:Master
order o_lvm inf: ms_drbd_data:promote cl_clvmd:start cl_lvm_cdata:start cl_fs_vm:start

The LVM clone is necessary so that you can deactivate the VG before disconnecting DRBD during a standby.  Not achieving this will STONITH the node.  The "--monitor y" change is absolutely necessary, or you won't even bring the VG online.  Starting clvmd inside Pacemaker might not be a necessary thing, but in this instance it seems to work very well.  It's also important to note that most of the init.d scripts related to this conundrum have been disabled: clvmd, drbd, to name two.

The GFS2 file system will not mount without gfs_controld running.  gfs_controld won't start on a clean Ubuntu Server 12.04 system because it seems the cman init script is fucked up.  Can't understand it, but inside /etc/init.d/cman you'll find a line that reads:
gfs_controld_enabled && cd /etc/init.d && ./gfs2-cluster start
Comment out this line and add this below it:
if [[ gfs_controld_enabled ]]; then
      cd /etc/init.d &&  ./gfs2-cluster start
fi
This will make the cman script actually CALL the gfs2-cluster script and thus start the gfs_controld daemon.  Shutdown seems to work correctly with no additional modifications.  You will find that once all these pieces are in place, GFS2 is viable on Ubuntu 12.04 AND you can bring your cluster up and down without watching your nodes commit creative suicide.

I honestly don't know why this is the way it is.  I wouldn't know where to even assign blame.  In the Ubuntu Server 12.04 Cluster Guide (work-in-progress), they suggest this resource:
primitive resGFSD ocf:pacemaker:controld \
        params daemon="gfs_controld" args="" \
        op monitor interval="120s"
This seems rather like a bastardization of what this resource agent is really for, but perhaps it works for them.  However, I would highly suspect this might suffer from the same issues that I ran into with OCFS2: that if CMAN isn't running the controld, putting a node into standby will wreak havoc on the node and cluster.  With OCFS2, the issue was in the ocfs2_controld daemon, which CMAN was all too happy to try to bring offline but would NOT under any circumstances that I could find start it up.

Once started by Pacemaker you also cannot seem to take it down, meaning the resource fails to stop and becomes a disqualifying offense for the node.   This issue seems unrelated to a missing killproc command that is non-standard among distributions, because even when you fix/fake it, the thing does not seem to accomplish anything.  ocfs2_controld continues to run in the background, and cman will fail to shutdown correctly after you try bringing a node down gracefully.  No ideas yet on how to fix this, but I might try for it next.  I had detailed making a working Ubuntu 12.04 OCFS2 cluster in a previous post...I will be double-checking those steps...

20121105

Useful

IPMI v2.0 - accessing SOL from Linux command line.

http://wiki.nikhef.nl/grid/Serial_Consoles

(use a bit rate that makes sense for you, only set if necessary)
ipmitool -I lanplus -H host.ipmi.nikhef.nl -U root sol set volatile-bit-rate 9.6
ipmitool -I lanplus -H host.ipmi.nikhef.nl -U root sol set non-volatile-bit-rate 9.6
 
ipmitool -I lanplus -H IPMI-BMC-IPADDR -U BMCPRIVUSER sol activate



20121102

Led Astray

It's frustrating, and it's my own damn fault.

I read in the HP v1910 switch documentation that an 802.3ad bond would utilize all connections for the transmission of data.   Even with static aggregation I thought I'd get something different than what, in fact, I received.  To quote their introduction on the concept of link aggregation:

"Link aggregation delivers the following benefits: * Increases bandwidth beyond the limits of any single link.  In an aggregate link, traffic is distributed across the member ports."
I'll spare you the rest.  It's my own damn fault because I took that little piece of marketing with an assumption:  That "traffic" indicated TCP packets regardless of their source or destination.  I know better now, and I do bow and scrape to the Prophets of Linux Bonding, the deities that espouse Whole Technical Truth.  I am not worthy!

Despite my best efforts, I cannot get more than 1G/sec between two LACP-connected machines.  Running iperf -s on one, and iperf -c on the other, the connection saturates as though a single channel were all that was available.  The only benefit then is that different machines are distributed across these multiple connections.  Those reading this and who knew better than I, I am sorry.  I'm an idiot.  May this blog serve to save others from my fate.

Static aggregation, as far as my HP switches are concerned, does nothing for mode-0 connections.  I can get a little better throughput, but watching the up-and-down of the flow rates suggests there is much evil happening, and I don't like it.  Plus, I can't really distribute a static aggregation across my switches as far as I know - maybe the HP switch stacking feature would help with this, but I also sense much evil there and don't want to go at it.

The only benefits I can derive from RR is by placing all connections into separate VLANs.  That, of course, kills any notion of redundancy and shared connectivity.  First, it's like having multiple switches, but if a single connection from a single machine goes down, then that whole machine is unable to communicate with the other machines across those virtual switches.  So, bollocks to that.

Second, it's damn hard to figure out a good, robust and non-impossible way to configure these VLANs to also communicate with the rest of the world.  I guess that it all boils down to my desire to use the maximum possible throughput to and from any given machine, without having to jump through hoops like creating gateway hosts just to aggregate all these connections into something recognizable by other networking hardware.  I am also not willing to sacrifice ports to the roles of active-passive, even though that would allow me at least one switch or link failure before catastrophic consequences took hold.

It's my own damn fault because I didn't take the time to read the bonding driver kernel documentation that the Good Lords of Kernel Development took the time to write.  I didn't, at least, until last night.  I poured through it, reading the telling tales of switches and support and the best way to get at certain kinds of redundancy or throughput.

802.3ad obviously doesn't do much for me either.  After reading the docs, I know this.  It does make aggregation on a single switch rather easy, but no more or less easy than mode-6 bonding.  Well, I take that back.  It IS less easy because the switch needs its ports configured.  It also doesn't support my need for multi-switch redundancy, so 802.3ad is out, too.

In short, if you're thinking of bonding two bonds together, don't.  It's just not worth it.  The trouble, the init scripts, the switch configuration will just not do you any good.  You'll still be stuck with 1 G/sec per machine connection.  Even worse, you might not even get your links quickly enough back if someone trips over the power-strip running your two highly-available switches.

I considered the VLAN solution, minus its connection to the world, thereby encapsulating my SAN-to-Hypervisor subnet in its own universe of ultra-high-throughput.  3 G/sec seemed a nice thing.  I managed to get close to that throughput; but, sadly, given that single-link failures would be catastrophic, I can't afford to take that risk.  Redundancy is too important.  I will relegate myself to mode-6, as it appears to be the most flexible, the most robust and the most reliable with regard to even link distribution.

I hope the price of 10GigE drops sooner rather than later...